Engineering deep dive · · 12 min read

Multi-Reward RL, Part 3: GDPO + CISPO + REPO-R at 27B, and the Advantage Floor That Stopped Learning

A controlled ladder on Qwen3.8-27B with a 12-channel verifier: swapping DAPO for CISPO, adding REPO-R and filtering prompts took the held-out score from −0.47 to +1.80. Visual explanations of each idea, real reward and holdout charts, and the advantage floor that silently zeroed 96% of the learning signal in a harsher environment.

Part 1 explained how PPO, GRPO, DAPO and GDPO turn several rewards into one learning signal. Part 2 benchmarked the trainers on a 14B model for 50 steps and found a trap: CISPO won the training curves and lost on held-out tasks. This part asks what happens when the same ideas get a bigger model and twelve times more steps.

We trained Qwen3.8-27B for 600 steps on a single-answer task scored by 12 reward channels, and climbed a controlled ladder: DAPO, then CISPO, then CISPO with REPO-R, then fewer but better prompts. Then we moved the winning recipe to a harsher environment and found a guard in our own advantage pipeline that silently threw away 96% of the learning signal.

The result in one line: same model, same rewards, same 600 steps. Held-out score went from −0.47 with DAPO to −0.02 with CISPO, to +1.33 with CISPO + REPO-R, and to +1.80 with a smaller, filtered prompt set. The starting model scores −1.37.

One step, four decisions1 · Score12 reward channelsper answer2 · GDPOnormalize each channelin its group, weight, sum3 · Guardno positive credit forinvalid or low-scoreanswers4 · REPO-Rreshape each token'sadvantage by its rarity5 · CISPOweight capped at 1.2,no token dropped6 · Updatetwo optimizer passesper rollout batchGDPO: what counts as better. Guard: who may be rewarded.REPO-R and CISPO: how hard each token moves.
The recipe as it runs in our trainer. Steps 2–5 are each a published method or a small guard we added; none of them changes the verifier or the rewards themselves.

1. Where Part 2 left us

On a 14B model and 50 steps, CISPO produced the highest training reward and one of the worst held-out scores. Adding REPO-R recovered a little (+0.0175), not enough to beat the untrained thinking baseline. Two questions stayed open: does the picture change with a larger model and a longer run, and which part of a combined recipe actually does the work?

So this time we changed one thing per run and judged every run the same way: a frozen held-out probe, and a checkpoint chosen without looking at it.

2. The setup, abstracted

We keep the domain out of this post on purpose; the reward structure is what matters. The policy reads a short specification (one of 500 prompts) and writes one structured answer of at most 128 tokens. A deterministic verifier parses it, runs it on a frozen dataset and returns twelve reward channels:

  • Core score: a cross-validated estimate of answer quality. This is the objective; GDPO priority 3.0.
  • Stability check: penalizes answers whose in-sample and out-of-sample ranks disagree; priority 1.0.
  • Action-cost penalty: every change in the answer’s output costs a fee.
  • Nine guards and shapers: a secondary predictiveness score, complexity, format, syntax and behavior novelty, a neutrality gate, a minimum-activity guard, a pattern guard and a stress test.

Held fixed in every run: Qwen3.8-27B in NF4 with a rank-32 adapter (alpha 64, FP32 master weights; LoRA unless a row says DoRA), no thinking, 4 prompts per step with 12 answers each, temperature 1.0, learning rate 1e-5, KL β 0.08, GDPO, 600 optimizer steps, one H100, seed 42. The REPO-R rows also change two loss settings, described in section 5.

Holdout: every 100 steps, a frozen probe of 50 unseen prompts is scored on a later slice of data that training never touched. Each run “ships” the checkpoint with the best train-side core score on the same probe, so the holdout number is never used to pick it.

3. GDPO in one minute

With twelve channels on very different scales, plain GRPO adds them first and normalizes the sum, so the widest channel decides almost everything. GDPO standardizes each channel inside the prompt group first, then adds them with explicit priorities. Every run in this post uses GDPO; Part 1 covers the math and an interactive demo.

Why GDPO normalizes each channel firstJoint (GRPO): sum, then normalizeGDPO: normalize each, then weightAABBCCDDcomplexity (spread 0.003)vanishes inside corecore × 3complexity × 1advantage (sum)Same answers, same rewards. With GDPO, B’s better complexity nearly closes its gap to A.
Illustration with made-up numbers. Part 1 has the interactive version and the math; in this study the GDPO priorities were core 3.0, stability 1.0, pattern guard 0.3, complexity 0.2 and 1.0 for the rest.

GDPO decides which answers count as better. The next two ideas decide how hard each token moves once we know.

4. CISPO: cap the weight, don’t drop the token

An answer was sampled by a slightly older copy of the policy. For each token, the ratio r = π_new / π_old says how much more likely the current policy makes it. PPO, GRPO and DAPO keep updates safe by clipping: once a token’s ratio leaves the allowed band, its gradient becomes zero for that step. CISPO (from the MiniMax-M1 report) keeps every token in the gradient and only caps its importance weight; we use a cap of 1.2.

DAPO: loss = −min(r·A, clip(r, 0.8, 1.28)·A) · CISPO: loss = −stopgrad(min(r, 1.2)) · A · log π(token)
DAPO / PPO-style clip (ε = 0.20 / 0.28)CISPO: weight capped at 1.2 (dashed)

Token in a better-than-average answer (A > 0)

0.00.51.01.50.60.811.21.41.61.8DAPO / PPO clipCISPO capratio r = π_new / π_old for one tokengradient weight

Past r = 1.28 the clipped DAPO token drops to 0.

Token in a worse-than-average answer (A < 0)

0.00.51.01.50.60.811.21.41.61.8DAPO / PPO clipCISPO capratio r = π_new / π_old for one tokengradient weight

Below r = 0.8 DAPO drops the token; above 1 it is not capped.

How hard one token is pushed, as a function of how much its probability already moved. The clipped objective switches the token off; CISPO keeps every token in the gradient and bounds only the size of its importance weight.

Why this matters for short structured answers: a handful of “fork” tokens decide what the answer does. Those are exactly the tokens whose probability moves fastest when a good answer is reinforced, so a clipped loss switches them off first. CISPO keeps teaching them, just more quietly.

Measured effect: swapping DAPO for CISPO, with everything else equal, moved the held-out score from −0.47 to −0.02, a gain of +0.45. In Part 2, CISPO overfit at 14B; here, with checkpoint selection by train-side core score and a 27B model, it generalized better than DAPO.

5. REPO-R: an entropy thermostat

RL tends to collapse a policy onto a few answers it already likes. Entropy, a measure of how many alternatives the policy still considers, falls, and exploration stops. REPO-R counters this per token: in a better-than-average answer, rare tokens get more credit; in a worse-than-average answer, rare tokens get less blame.

A > 0: A′ = A · (1 − ζ · log p) · A < 0: A′ = A · (1 + ζ · log p) · where p is the token’s probability and log p < 0
token in a better answer, ζ = 0.05token in a worse answer, ζ = 0.05either, ζ = 0.0008 (our median) (dashed)
×0.6×0.8×1.0×1.2×1.4-3-2.5-2-1.5-1-0.50A > 0A < 0median ζ, A > 0median ζ, A < 0token probability p (log10): rare ← → commonadvantage multiplier
REPO-R in one picture: rare tokens get more credit when the answer was good and less blame when it was bad, so the policy keeps its alternatives. At ζ = 0.05 a 1-in-1000 token gets ×1.35 credit or ×0.65 blame; at our median ζ the effect is under 1%.

The strength ζ is not a fixed knob. A controller records the average entropy of the first five updates as its target (the “w5” in our run names). When entropy falls below target, ζ doubles; when it rises above, ζ halves, and it never goes below zero. Here is the real log of our best run:

policy entropy (answer tokens)target 0.542 = mean of the first 5 updates (dashed)ζ used (right panel)
0.30.40.50.60.70.80200400600entropytargetoptimizer stepentropy0.000.010.020.030.040.050200400600zetaoptimizer stepζ
Real controller log of our best run (1790958861). REPO-R mostly sleeps: median ζ = 0.00078, above 0.01 on 27% of steps. When entropy dips below target, ζ doubles each update up to its 0.05 cap; when entropy recovers, it halves back toward zero.

Measured effect: adding REPO-R to CISPO moved held-out from −0.02 to +1.33, a gain of +1.35, the largest single step in the study. A caveat we take seriously: the REPO-R configuration also switches on full-token loss (every token enters the loss, not only the 20% with the highest entropy) and two optimizer passes per rollout batch. With ζ this small most of the time, part of the gain may come from those two changes. The ablation that would separate them is still on our list.

6. Results at 27B

FINDING 01

CISPO beat DAPO

−0.47 → −0.02

One-line loss change, everything else fixed. Part 2’s 14B, 50-step benchmark had favored DAPO.

FINDING 02

REPO-R bundle: biggest step

−0.02 → +1.33

Includes full-token loss and two passes per batch, which we have not yet isolated.

FINDING 03

Fewer, better prompts

500 prompts +1.33 · 226 prompts +1.80

Keeping only prompts where the starting model sometimes, but not always, succeeds gave the best run.

FINDING 04

Same curves, 2× holdout gap

DAPO +0.69 vs CISPO +1.33

With REPO-R and DoRA on both, the training curves overlap. Only the holdout separates them.

DAPO · LoRA · 500 promptsDAPO · DoRA · 500 promptsCISPO · LoRA · 500 promptsCISPO + REPO-R · LoRA · 500 promptsCISPO + REPO-R · LoRA · 226 promptsCISPO + REPO-R · DoRA · 226 promptsDAPO + REPO-R · DoRA · 226 prompts
-1.5-1.0-0.50.00.51.01.52.00100200300400500600DAPO · LoRA · 500 promptsDAPO · DoRA · 500 promptsCISPO · LoRA · 500 promptsCISPO + REPO-R · LoRA · 500 promptsCISPO + REPO-R · LoRA · 226 promptsCISPO + REPO-R · DoRA · 226 promptsDAPO + REPO-R · DoRA · 226 promptsDAPO · LoRA · 500 prompts: -0.474 at step 600DAPO · DoRA · 500 prompts: -0.287 at step 200CISPO · LoRA · 500 prompts: -0.024 at step 600CISPO + REPO-R · LoRA · 500 prompts: 1.327 at step 600CISPO + REPO-R · LoRA · 226 prompts: 1.800 at step 500CISPO + REPO-R · DoRA · 226 prompts: 1.333 at step 500DAPO + REPO-R · DoRA · 226 prompts: 0.694 at step 500checkpoint step (0 = starting model)holdout score
Frozen 50-prompt probe on an evaluation slice that training never saw, scored every 100 steps. Stars mark the checkpoint each run would ship, chosen by its train-side core score, not by holdout. One seed per configuration.
Environment A, Qwen3.8-27B, 600 steps, one seed each. Holdout = frozen 50-prompt probe on an unseen evaluation slice, at the checkpoint chosen by train-side core score. “Change” compares with the row named in “versus”.
Step of the ladderConfigurationShipped stepHoldoutChangeVersus
Start: DAPO lossDAPO · LoRA · 500 prompts600−0.47—starting model: −1.37
Swap the loss for CISPOCISPO · LoRA · 500 prompts600−0.02+0.45DAPO · LoRA · 500 prompts
Add REPO-R (with full-token loss, 2 passes)CISPO + REPO-R · LoRA · 500 prompts600+1.33+1.35CISPO · LoRA · 500 prompts
Train on 226 selected prompts instead of 500best holdoutCISPO + REPO-R · LoRA · 226 prompts500+1.80+0.47CISPO + REPO-R · LoRA · 500 prompts
Swap LoRA for DoRACISPO + REPO-R · DoRA · 226 prompts500+1.33−0.47CISPO + REPO-R · LoRA · 226 prompts
Swap CISPO back to DAPO (keeps REPO-R, DoRA)DAPO + REPO-R · DoRA · 226 prompts500+0.69−0.64CISPO + REPO-R · DoRA · 226 prompts

DoRA did not help once REPO-R was on: +1.33 against +1.80 for the same recipe with LoRA. Without REPO-R it gave DAPO a small lift (−0.29 against −0.47). The best run peaked at step 500 and sagged slightly by step 600, which is why we ship by the train-side selector rather than the last step.

Training curves are not a verdict

CISPO + REPO-R · DoRA · 226 promptsDAPO + REPO-R · DoRA · 226 prompts
-1.0-0.50.00.51.00200400600CISPO + REPO-R · DoRA · 226 promptsDAPO + REPO-R · DoRA · 226 promptsoptimizer steptrain core (10-step mean)-1.5-1.0-0.50.00.51.01.50200400600CISPO + REPO-R · DoRA · 226 promptsDAPO + REPO-R · DoRA · 226 promptsCISPO + REPO-R · DoRA · 226 prompts: 1.333 at step 500DAPO + REPO-R · DoRA · 226 prompts: 0.694 at step 500checkpoint stepholdout score
Same adapter, prompts and REPO-R; only the loss differs. The training curves are hard to tell apart, the held-out scores are not: 1.33 for CISPO, 0.69 for DAPO.

This is Part 2’s lesson again, at a larger scale: a trainer can look identical on the training reward and still learn something quite different. Every decision in this post is made on a frozen holdout, and every run uses the same probe and the same selection rule.

Every training curve

The holdout chart is what we decide on; the curves below are what the trainer saw. Switch between the two environments and five logged metrics, and hide runs to compare any subset. In Environment A, notice how close the REPO-R runs stay on training reward and core score while their held-out scores spread from +0.69 to +1.80.

-4-3-2-10120100200300400500600DAPO · LoRA · 500 promptsDAPO · DoRA · 500 promptsCISPO · LoRA · 500 promptsCISPO + REPO-R · LoRA · 500 promptsCISPO + REPO-R · LoRA · 226 promptsCISPO + REPO-R · DoRA · 226 promptsDAPO + REPO-R · DoRA · 226 promptsoptimizer steptrain reward
Every run in this post, as logged by the trainer: 10-step means, one point per 5 logged steps (REPO-R runs log once per rollout batch, every second optimizer step). Environment B no-floor runs were still training when we took this snapshot, so they end early. Tick a run to hide or show it; the full per-step values are in the JSON download.

7. The advantage floor trap

Next we moved the winning recipe to Environment B: different data, the same verifier, and per-action costs about 19× higher. Same 27B model, same GDPO, CISPO and REPO-R, and the same prompt filter, which this time found so few prompts with any successful answer that it kept the 100 closest to break-even. The stress-test channel is off there: its world model failed its own quality gates on that data. The runs started, the curves wiggled, and almost nothing happened.

The cause was a guard we had added to GDPO ourselves. It keeps invalid answers, and answers whose core score is at or below a floor (0), from receiving positive advantage, while preserving the rule that advantages inside a group sum to zero. In Environment A this is a sensible safety net: good answers score above zero. In Environment B, the fresh policy almost never does, and only about 1% of answers clear the floor. A group in which all twelve answers are below the floor gets twelve zeros.

The floor trap: twelve answers, zero signalCore score of 12 answers to one promptfloor = 0-6-4-2All 12 are below the floor, so all 12may receive only non-positive advantage,and the group must still sum to zero.Floor at 012 × 0.00 → no gradient at allNo floor (only invalid answers bounded)less-bad answers still win
Exact output of our projection code on this group; the sixth answer is invalid, so it stays negative either way. The guard was written for an environment where good answers score above zero. In a harsher one, every answer of a fresh policy scores below zero, and the guard turns whole groups into silence.

The logs made the damage plain. With the floor, 4% of answers carried any learning signal; KL to the starting model stayed around 0.02. Removing the floor, and changing nothing else, gave 100%, a KL rising to about 0.2, and a training reward that climbed from about −10 to −2.5.

advantage floor at 0 (run 1791045171) (dashed)no floor (run 1791130455)
0%25%50%75%100%0200400floor at 0no floorstepsignal share0.00.10.20.30200400floor at 0no floorstepKL to start-15-10-500200400floor at 0no floorsteptrain reward
Environment B, same prompts and recipe, one change: the advantage floor. Ten-step means; the no-floor run was still training when we took this snapshot (339 of 600 steps). A second data slice repeated the pattern: 3.5% vs 99.8% non-zero advantages.

All seven Environment B runs, including two more data slices, are in the curve explorer in section 6: every floor-at-zero run stays flat, every no-floor run climbs.

Turning the guard off did not turn it off

Our first fix was to disable the guard. The trainer then falls back to a second safety net that clamps any positive advantage to zero when the core score is at or below zero: the same trap with a different name. What worked was keeping the projection, which still forces invalid answers to non-positive advantage, and moving the floor far below the environment’s score range. A floor is an assumption about the reward distribution. When the environment changes, the assumption has to be checked again.

Even in Environment A the floor was not free: 10–41% of advantages were zero across runs, more in the runs that learned slowly. Whether a lower floor would also help there is one of our next experiments. The held-out results for Environment B will follow once those runs finish their probes.

8. What we would do again

  1. Log the share of non-zero advantages from step 1. In a healthy run it sits well above half. At 4% the trainer is idling while the GPU bill keeps running.
  2. Treat every guard threshold as a property of the environment. Re-derive floors and clamps when costs, data or reward ranges change, and check what the fallback path does when you disable one.
  3. Change one thing per run and judge on a frozen holdout. Ship the checkpoint chosen by a train-side selector, and use the holdout only to verify it.
  4. Prefer fewer, informative prompts. Prompts the starting model always or never solves give little signal; filtering them beat training on all 500.
  5. Name the recipe by its parts. GDPO + CISPO + REPO-R is a combination of published methods plus one guard of ours, and it should be credited that way.

Limits: one seed per configuration; one held-out slice; REPO-R measured as a bundle with full-token loss and two passes; Environment B holdout pending. Treat the gaps as strong signals, not final effect sizes.

Evidence and method references

The JSON export holds every logged optimizer step (reward, core score, cost, KL, entropy), the share of non-zero advantages per step, the REPO-R controller log and the holdout score of every checkpoint, with run IDs. The Markdown companion has the tables.

Methods: DeepSeekMath / GRPO; DAPO; GDPO; MiniMax-M1 / CISPO; REPO-R.