Part 1 explained how PPO, GRPO, DAPO and GDPO turn several rewards into one learning signal. Part 2 benchmarked the trainers on a 14B model for 50 steps and found a trap: CISPO won the training curves and lost on held-out tasks. This part asks what happens when the same ideas get a bigger model and twelve times more steps.
We trained Qwen3.8-27B for 600 steps on a single-answer task scored by 12 reward channels, and climbed a controlled ladder: DAPO, then CISPO, then CISPO with REPO-R, then fewer but better prompts. Then we moved the winning recipe to a harsher environment and found a guard in our own advantage pipeline that silently threw away 96% of the learning signal.
The result in one line: same model, same rewards, same 600 steps. Held-out score went from −0.47 with DAPO to −0.02 with CISPO, to +1.33 with CISPO + REPO-R, and to +1.80 with a smaller, filtered prompt set. The starting model scores −1.37.
1. Where Part 2 left us
On a 14B model and 50 steps, CISPO produced the highest training reward and one of the worst held-out scores. Adding REPO-R recovered a little (+0.0175), not enough to beat the untrained thinking baseline. Two questions stayed open: does the picture change with a larger model and a longer run, and which part of a combined recipe actually does the work?
So this time we changed one thing per run and judged every run the same way: a frozen held-out probe, and a checkpoint chosen without looking at it.
2. The setup, abstracted
We keep the domain out of this post on purpose; the reward structure is what matters. The policy reads a short specification (one of 500 prompts) and writes one structured answer of at most 128 tokens. A deterministic verifier parses it, runs it on a frozen dataset and returns twelve reward channels:
- Core score: a cross-validated estimate of answer quality. This is the objective; GDPO priority 3.0.
- Stability check: penalizes answers whose in-sample and out-of-sample ranks disagree; priority 1.0.
- Action-cost penalty: every change in the answer’s output costs a fee.
- Nine guards and shapers: a secondary predictiveness score, complexity, format, syntax and behavior novelty, a neutrality gate, a minimum-activity guard, a pattern guard and a stress test.
Held fixed in every run: Qwen3.8-27B in NF4 with a rank-32 adapter (alpha 64, FP32 master weights; LoRA unless a row says DoRA), no thinking, 4 prompts per step with 12 answers each, temperature 1.0, learning rate 1e-5, KL β 0.08, GDPO, 600 optimizer steps, one H100, seed 42. The REPO-R rows also change two loss settings, described in section 5.
Holdout: every 100 steps, a frozen probe of 50 unseen prompts is scored on a later slice of data that training never touched. Each run “ships” the checkpoint with the best train-side core score on the same probe, so the holdout number is never used to pick it.
3. GDPO in one minute
With twelve channels on very different scales, plain GRPO adds them first and normalizes the sum, so the widest channel decides almost everything. GDPO standardizes each channel inside the prompt group first, then adds them with explicit priorities. Every run in this post uses GDPO; Part 1 covers the math and an interactive demo.
GDPO decides which answers count as better. The next two ideas decide how hard each token moves once we know.
4. CISPO: cap the weight, don’t drop the token
An answer was sampled by a slightly older copy of the policy. For each token, the ratio r = π_new / π_old says how much more likely the current policy makes it. PPO, GRPO and DAPO keep updates safe by clipping: once a token’s ratio leaves the allowed band, its gradient becomes zero for that step. CISPO (from the MiniMax-M1 report) keeps every token in the gradient and only caps its importance weight; we use a cap of 1.2.
Token in a better-than-average answer (A > 0)
Past r = 1.28 the clipped DAPO token drops to 0.
Token in a worse-than-average answer (A < 0)
Below r = 0.8 DAPO drops the token; above 1 it is not capped.
Why this matters for short structured answers: a handful of “fork” tokens decide what the answer does. Those are exactly the tokens whose probability moves fastest when a good answer is reinforced, so a clipped loss switches them off first. CISPO keeps teaching them, just more quietly.
Measured effect: swapping DAPO for CISPO, with everything else equal, moved the held-out score from −0.47 to −0.02, a gain of +0.45. In Part 2, CISPO overfit at 14B; here, with checkpoint selection by train-side core score and a 27B model, it generalized better than DAPO.
5. REPO-R: an entropy thermostat
RL tends to collapse a policy onto a few answers it already likes. Entropy, a measure of how many alternatives the policy still considers, falls, and exploration stops. REPO-R counters this per token: in a better-than-average answer, rare tokens get more credit; in a worse-than-average answer, rare tokens get less blame.
The strength ζ is not a fixed knob. A controller records the average entropy of the first five updates as its target (the “w5” in our run names). When entropy falls below target, ζ doubles; when it rises above, ζ halves, and it never goes below zero. Here is the real log of our best run:
Measured effect: adding REPO-R to CISPO moved held-out from −0.02 to +1.33, a gain of +1.35, the largest single step in the study. A caveat we take seriously: the REPO-R configuration also switches on full-token loss (every token enters the loss, not only the 20% with the highest entropy) and two optimizer passes per rollout batch. With ζ this small most of the time, part of the gain may come from those two changes. The ablation that would separate them is still on our list.
6. Results at 27B
CISPO beat DAPO
−0.47 → −0.02One-line loss change, everything else fixed. Part 2’s 14B, 50-step benchmark had favored DAPO.
REPO-R bundle: biggest step
−0.02 → +1.33Includes full-token loss and two passes per batch, which we have not yet isolated.
Fewer, better prompts
500 prompts +1.33 · 226 prompts +1.80Keeping only prompts where the starting model sometimes, but not always, succeeds gave the best run.
Same curves, 2× holdout gap
DAPO +0.69 vs CISPO +1.33With REPO-R and DoRA on both, the training curves overlap. Only the holdout separates them.
| Step of the ladder | Configuration | Shipped step | Holdout | Change | Versus |
|---|---|---|---|---|---|
| Start: DAPO loss | DAPO · LoRA · 500 prompts | 600 | −0.47 | — | starting model: −1.37 |
| Swap the loss for CISPO | CISPO · LoRA · 500 prompts | 600 | −0.02 | +0.45 | DAPO · LoRA · 500 prompts |
| Add REPO-R (with full-token loss, 2 passes) | CISPO + REPO-R · LoRA · 500 prompts | 600 | +1.33 | +1.35 | CISPO · LoRA · 500 prompts |
| Train on 226 selected prompts instead of 500best holdout | CISPO + REPO-R · LoRA · 226 prompts | 500 | +1.80 | +0.47 | CISPO + REPO-R · LoRA · 500 prompts |
| Swap LoRA for DoRA | CISPO + REPO-R · DoRA · 226 prompts | 500 | +1.33 | −0.47 | CISPO + REPO-R · LoRA · 226 prompts |
| Swap CISPO back to DAPO (keeps REPO-R, DoRA) | DAPO + REPO-R · DoRA · 226 prompts | 500 | +0.69 | −0.64 | CISPO + REPO-R · DoRA · 226 prompts |
DoRA did not help once REPO-R was on: +1.33 against +1.80 for the same recipe with LoRA. Without REPO-R it gave DAPO a small lift (−0.29 against −0.47). The best run peaked at step 500 and sagged slightly by step 600, which is why we ship by the train-side selector rather than the last step.
Training curves are not a verdict
This is Part 2’s lesson again, at a larger scale: a trainer can look identical on the training reward and still learn something quite different. Every decision in this post is made on a frozen holdout, and every run uses the same probe and the same selection rule.
Every training curve
The holdout chart is what we decide on; the curves below are what the trainer saw. Switch between the two environments and five logged metrics, and hide runs to compare any subset. In Environment A, notice how close the REPO-R runs stay on training reward and core score while their held-out scores spread from +0.69 to +1.80.
7. The advantage floor trap
Next we moved the winning recipe to Environment B: different data, the same verifier, and per-action costs about 19× higher. Same 27B model, same GDPO, CISPO and REPO-R, and the same prompt filter, which this time found so few prompts with any successful answer that it kept the 100 closest to break-even. The stress-test channel is off there: its world model failed its own quality gates on that data. The runs started, the curves wiggled, and almost nothing happened.
The cause was a guard we had added to GDPO ourselves. It keeps invalid answers, and answers whose core score is at or below a floor (0), from receiving positive advantage, while preserving the rule that advantages inside a group sum to zero. In Environment A this is a sensible safety net: good answers score above zero. In Environment B, the fresh policy almost never does, and only about 1% of answers clear the floor. A group in which all twelve answers are below the floor gets twelve zeros.
The logs made the damage plain. With the floor, 4% of answers carried any learning signal; KL to the starting model stayed around 0.02. Removing the floor, and changing nothing else, gave 100%, a KL rising to about 0.2, and a training reward that climbed from about −10 to −2.5.
All seven Environment B runs, including two more data slices, are in the curve explorer in section 6: every floor-at-zero run stays flat, every no-floor run climbs.
Turning the guard off did not turn it off
Our first fix was to disable the guard. The trainer then falls back to a second safety net that clamps any positive advantage to zero when the core score is at or below zero: the same trap with a different name. What worked was keeping the projection, which still forces invalid answers to non-positive advantage, and moving the floor far below the environment’s score range. A floor is an assumption about the reward distribution. When the environment changes, the assumption has to be checked again.
Even in Environment A the floor was not free: 10–41% of advantages were zero across runs, more in the runs that learned slowly. Whether a lower floor would also help there is one of our next experiments. The held-out results for Environment B will follow once those runs finish their probes.
8. What we would do again
- Log the share of non-zero advantages from step 1. In a healthy run it sits well above half. At 4% the trainer is idling while the GPU bill keeps running.
- Treat every guard threshold as a property of the environment. Re-derive floors and clamps when costs, data or reward ranges change, and check what the fallback path does when you disable one.
- Change one thing per run and judge on a frozen holdout. Ship the checkpoint chosen by a train-side selector, and use the holdout only to verify it.
- Prefer fewer, informative prompts. Prompts the starting model always or never solves give little signal; filtering them beat training on all 500.
- Name the recipe by its parts. GDPO + CISPO + REPO-R is a combination of published methods plus one guard of ours, and it should be credited that way.
Limits: one seed per configuration; one held-out slice; REPO-R measured as a bundle with full-token loss and two passes; Environment B holdout pending. Treat the gaps as strong signals, not final effect sizes.
Evidence and method references
The JSON export holds every logged optimizer step (reward, core score, cost, KL, entropy), the share of non-zero advantages per step, the REPO-R controller log and the holdout score of every checkpoint, with run IDs. The Markdown companion has the tables.
Methods: DeepSeekMath / GRPO; DAPO; GDPO; MiniMax-M1 / CISPO; REPO-R.