We increased the compute spent on reinforcement learning (RL) post-training by 10x, starting from DeepSeek V4 Pro, a 1.6 trillion parameter base model. Across the run, the model kept improving as compute went up. We did not see the plateau that is often assumed for RL post-training: sustained RL continued to pay off, and the scaling laws held.
This post explains the three policy-optimization methods behind that run (GRPO, DAPO and GSPO), what each one changes and why, and the engineering problems that come with running RL on a model this large.
RL post-training in one paragraph
A pretrained model already has broad capabilities. RL post-training shapes which of those behaviours it actually uses. For each prompt, the current model (the policy) samples responses, a reward function scores them, and the model is updated to make high-reward responses more likely and low-reward ones less likely. Most of the difficulty is in doing that update in a way that stays stable over thousands of steps, and in generating enough samples fast enough to feed it.
GRPO: Group Relative Policy Optimization
GRPO was introduced by DeepSeek as a simplification of PPO. PPO estimates how good each response is relative to expectations using a learned value model (a critic), usually about as large as the policy itself. At 1.6T parameters, a second model of that size roughly doubles the memory and compute bill.
GRPO removes the critic. For each prompt q, it samples a group of G responses from the current policy, scores each one, and uses the group itself as the baseline:
Every token in response i receives the same advantage Â_i. A response that beats its siblings is pushed up; one that trails them is pushed down. The update itself is PPO's clipped objective, applied per token:
The ratio r_i,t measures how much the policy has moved since the responses were sampled, and clipping stops any single update from moving it too far. The KL term keeps the policy close to a frozen reference model.
GRPO works well, but two details matter at scale. First, the importance ratio is computed per token from a single sample, which makes it a noisy correction that accumulates over long responses. Second, the 1/|o_i| averaging gives every response equal weight, so each token in a long response counts for less than a token in a short one. DAPO and GSPO each address one of these.
DAPO: Decoupled Clip and Dynamic Sampling
DAPO keeps GRPO's group-relative advantage and adds four practical changes aimed at long, stable runs:
- Clip-Higher. PPO clips the ratio symmetrically to [1−ε, 1+ε]. DAPO decouples the two bounds and raises the upper one (ε_high > ε_low). A symmetric clip makes it hard for low-probability tokens to grow, so exploration shrinks and entropy collapses; a looser upper bound lets promising but unlikely tokens become more likely.
- Dynamic sampling. If every response in a group gets the same reward (all correct or all wrong), every advantage is zero and the group contributes no gradient. DAPO over-samples prompts and drops those groups, so each batch is full of samples that actually teach the model something.
- Token-level loss. Instead of averaging within each response and then across responses, DAPO averages over every token in the batch. Long responses contribute in proportion to their length, so both good long reasoning and bad long patterns (repetition, rambling) are weighted properly.
- Overlong reward shaping. Responses cut off at the length limit are masked out of the loss rather than scored as failures, and a soft penalty applies as responses approach the limit. This keeps length from drifting upwards without adding noise from truncation.
DAPO also drops the KL penalty, on the grounds that during long RL runs the policy is supposed to move well away from its starting point.
GSPO: Group Sequence Policy Optimization
GSPO, from the Qwen team, targets the first problem: token-level importance ratios. Rewards are given to whole responses, but GRPO corrects for policy drift token by token. GSPO makes the unit of optimization match the unit of reward by defining one importance ratio per sequence:
s_i is the geometric mean of the per-token ratios, so it is normalized for length and comparable across short and long responses. Clipping now keeps or drops whole responses instead of individual tokens, and every token in a response is weighted by the same factor. Because the ratio is an average, it stays much closer to 1, so the clipping range is far narrower in absolute terms than PPO's usual 0.2.
This matters most for Mixture-of-Experts models like the DeepSeek family. After an update, a small change in the router can send a token to different experts, so that token's probability can jump even though the model as a whole barely moved. Token-level ratios see those jumps as large policy changes and destabilize training. A sequence-level ratio averages them out. The GSPO authors report that this stabilizes MoE RL training without workarounds such as replaying the old routing decisions.
How the three fit together
These methods are complementary rather than competing. GRPO provides the group-relative advantage that both others build on and removes the critic. GSPO changes how the policy's movement is measured and constrained (per sequence rather than per token). DAPO changes what goes into each batch and how it is weighted: which prompts are kept, how tokens are counted, how length is handled, and how much room there is to explore.
Challenges of RL at 1.6T parameters
Memory
In BF16, the weights of a 1.6T parameter model take about 3.2 TB. Full training state (BF16 weights and gradients plus FP32 master weights and two optimizer moments, around 16 bytes per parameter) is roughly 25 TB. None of this fits on one node, so the model is sharded across hundreds of GPUs with a combination of data, tensor, pipeline and expert parallelism. Dropping the critic (GRPO) is what keeps the total within reach.
Generation dominates the step
Every RL step starts by sampling many responses per prompt, and autoregressive generation is slow compared with a training pass. Rollouts run on a dedicated high-throughput inference engine such as vLLM or SGLang, alongside the trainer. After every update, the new weights (terabytes of them) have to be pushed to the inference workers before the next batch of rollouts. Weight synchronization becomes a first-class systems problem.
Long tails and off-policy data
Response lengths vary a lot, and a synchronous step waits for its slowest sample. Keeping GPUs busy means overlapping generation with training, which makes some data slightly stale: it was sampled by a policy a step or two old. Importance ratios exist to correct for exactly this, which is one more reason their stability (GSPO) matters.
Training–inference mismatch
The inference engine and the trainer compute token probabilities with different kernels, precisions and batch shapes, so their log-probabilities never match exactly. In an MoE model, even routing can differ between the two. To the optimizer, this looks like off-policy noise. It has to be measured continuously, for example by tracking the KL between the rollout engine's log-probabilities and the trainer's, and kept small.
Keeping a long run healthy
Sustained improvement over a 10x longer run only happens if the usual failure modes stay under control: entropy collapse, responses growing without bound, and the policy learning to exploit the reward rather than improve. Entropy, response length, clip fraction and held-out evaluations are all monitored throughout; Clip-Higher and overlong shaping (DAPO) exist precisely to keep the first two in check.
Operations
At this scale, hardware faults during a long run are expected rather than exceptional. Frequent sharded checkpoints and automatic restarts are what turn a long run from fragile into routine.
What's next
The main result is simple: we have not yet found the point where more RL compute stops helping. We will keep scaling it.
If you want to work on this, we are hiring an AI Researcher: Post-Training and engineers for AI infrastructure and GPU orchestration.
References
- Shao et al., DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (introduces GRPO), 2024.
- Yu et al., DAPO: An Open-Source LLM Reinforcement Learning System at Scale, 2025.
- Zheng et al., Group Sequence Policy Optimization, 2025.