Kurisutina

Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model?

Yang Yue (project lead), Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, Gao Huang; LeapLab, Tsinghua University and Shanghai Jiao Tong University. arXiv:2504.13837v5 [cs.AI], 24 November 2025. Project page: limit-of-RLVR.github.io. Read by researcher R6 (capability), 6 October 2026, as the test of "the slot can shape but not add capacity the base lacks". Provenance: papers/capability/yue2025_rl_reasoning_capacity.provenance.json.

What was read

  • Read in full: the pdftotext conversion (3,845 lines, 31 pages): sections 1–7, author contributions, references, and Appendices A–E.
    • A: algorithms, the pass@k estimator;
    • B: related work;
    • C: more results, chain-of-thought validity, accuracy histograms, perplexity, algorithms, KL and rollouts, coverage Tables 2–6, temperature and entropy, training dynamics, worked chain-of-thought cases;
    • D: prompts;
    • E: impacts.
  • Not read: pass@k curves and histograms as images. Lines holding only histogram counts were passed over; Tables 2–6 values were read from the text layer.

What they did

  • Question: does reinforcement learning with verifiable rewards (RLVR) give a language model reasoning abilities its base model lacks, or only make it sample existing abilities more reliably?
  • Metric: pass@k at large k (up to 1,024), the share of problems solved by any of k samples, as the model's "reasoning boundary". It uses an unbiased estimator. Chains of thought were checked by hand on the hardest problems.
  • Models and tasks:
    • math with base models: Qwen2.5-7B/14B/32B and LLaMA-3.1-8B, with SimpleRLZoo GRPO models, Oat-Zero-7B and DAPO-32B;
    • code: Code-R1 and DeepCoder-14B;
    • visual math: Qwen2.5-VL-7B;
    • a near-frontier pair: Mistral-Medium-3 against Magistral-Medium.
  • Controlled training of six RL algorithms (PPO, GRPO, Reinforce++, RLOO, ReMax, DAPO) on Omni-MATH from Qwen2.5-7B.
  • Comparison with distillation: DeepSeek-R1-Distill-Qwen-7B against its base, Qwen2.5-Math-7B.

Main results (verified)

  • RL raises pass@1 but narrows the boundary.

    • In every benchmark and family, RL models win at small k, and base models catch up and overtake at large k.
    • Example: Qwen2.5-32B base beats its RL model by about 9% on Minerva at k = 128.
  • What RL solves, the base already could (Table 2):

    AIME24 (k = 1024) MATH500 (k = 128)
    Both solve 63.3% 92.4%
    Base only 13.3% 3.6%
    RL only 0.0% 1.0%
    Neither 23.3% 3.0%

    Even the 1% "RL only" problems on MATH500 were solved by the base at 1,024 samples.

  • More training makes the trade sharper (Table 4, Qwen2.5-7B with GRPO):

    • Training-set pass@1 rose from 9.9 to 42.5 over 450 steps.
    • In-domain pass@256 fell from 69.1 to 63.9; MATH500 pass@256 from 96.2 to 95.4.
  • RL outputs are already likely under the base model. The base model's perplexity on RL outputs matches its perplexity on its own likely outputs, and falls as training proceeds. RL "mainly sharpens the distribution within the base model's prior".

  • Algorithms barely differ.

    • Across six algorithms, MATH500 pass@1 is 73.5–75.6 (base 34.5) and in-domain pass@256 is 67.0–69.7 (base 69.1).
    • The gap between RL pass@1 and the base's pass@256 stays above 40 points, "far from optimal".
  • Entropy and KL are not the explanation. Matching the RL model's output entropy to the base's (higher temperature) did not restore its coverage. A KL penalty kept pass@1 and lowered pass@128. More rollouts per prompt (32 against 8) helped slightly.

  • Distillation is different. The distilled model's pass@k curve is "consistently and significantly above" the base's at every k. Distillation "introduces new reasoning patterns learned from a stronger teacher".

  • The pattern holds near the frontier. Magistral-Medium solves about 7 more AIME24 and 8 more AIME25 problems at k = 1, with the gap closing as k grows.

  • The authors' explanation:

    • The pretrained prior guides exploration in a huge action space.
    • Samples outside the prior are mostly invalid, so policy gradients reinforce in-prior paths.
    • The prior is "a double-edged sword". Proposed remedies: exploration at higher abstraction, curricula, process rewards, multi-turn agentic RL.

Limits

  • Only current RLVR recipes, at RL compute far below pretraining scale. Whether RL at pretraining-scale compute breaks out of the boundary is open (the authors say so).
  • pass@k at large k can count lucky guesses in math. Hand checks found valid chains of thought for most hardest-solved problems: base 24/25 on GSM8K, 5/6 on AIME24.
  • Most of the best RL systems are proprietary. Code and vision used instruction-tuned starting points, not bases.

What it means for Kurisutina (inference)

  • "The slot can lower and shape capability but cannot add capacity the base lacks" is half right.
    • Reward-driven training of a slot on outcomes sharpens what the base can already do and narrows the rest.
    • Imitation of a stronger source (distillation) does add reasoning patterns the base lacked.
    • Amadeus's slot is trained by imitating a person's behaviour. If the person reasons in ways the base cannot produce, the slot can import those patterns, up to the slot's own capacity and the base's architecture (depth, compute per token).
  • What the base bounds is architectural capacity, not every pattern. The base must be able to represent the person's computation (depth and width, as in the looped-transformer papers). The patterns themselves can come in through the slot.
  • Training choice matters for fidelity. Outcome-reward fine-tuning would push a replica towards higher pass@1 on what it already can do, and towards less diversity. That is overshoot plus homogenisation, the opposite of reproducing a person's spread of attempts and errors. Imitation objectives fit a replica better.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.