Rui-Jie Zhu, Zixuan Wang, Kai Hua, Tianyu Zhang, Ziniu Li, Haoran Que and 30 others, with Yoshua Bengio and Jason Eshraghian. ByteDance Seed, UC Santa Cruz, Princeton, Mila and others. arXiv:2510.25741v5 [cs.CL], 1 July 2026 (first version October 2025; preprint). Models released at ouro-llm.github.io. Read by researcher R6 (capability), 6 October 2026, as the LoopLM that MELT is converted from. Provenance: papers/capability/zhu2025_ouro_looped_lm.provenance.json.
What was read
- Read in full: the pdftotext conversion (8,290 lines, 54 pages): abstract, sections 1–8, contributions, references, and Appendices A–E.
- A: choice of prior;
- B: knowledge capacity, Mano and multi-hop QA tasks, MMLU category table, theory and proof;
- C: evaluation settings;
- D: small-scale scaling study;
- E: scaling-law generalisation.
- Figures were not read as images. In Figures 13–23 (scaling-law panels), lines holding only axis ticks or panel labels were filtered out; every caption and all prose were read.
- Two numbers are read from figure text layers, with the mapping inferred: Figure 1's radar values and Figure 6's Mano bars. They are marked where used.
- Tables 7–9 are scrambled by pdftotext. Their values were matched to models through the paper's own text, e.g. "BBH 71.02 vs 70.95".
What they did
-
Architecture (LoopLM). One stack of N decoder layers is applied T times per token, with the same weights every time.
- Ouro 1.4B: 24 layers, width 2048, multi-head attention (no GQA), SwiGLU, RoPE, sandwich RMSNorm, vocabulary 49,152.
- Ouro 2.6B: 48 layers, made by duplicating the 1.4B's 24 layers ("upcycling").
- Both run T = 4 loops. An LM head and an exit gate act after every loop.
-
Adaptive depth. The objective is the expected loss over exit steps minus β × the entropy of the exit distribution, an ELBO with a uniform prior. A second stage trains only the gate, to exit when the next loop's loss improvement falls below a threshold. At inference, exit when the cumulative exit probability crosses q (the "Q-exit" rule).
-
Training, 7.7T tokens in all, open data only:
Stage Tokens Loops Notes 1a, pre-training I 3T 8 Shared by both models 1b, pre-training II 3T 4 The 2.6B is upcycled here 2, annealing 1.4T 4 Higher-quality data 3, long context 20B 4 64K-token sequences 4, mid-training 300B 4 90B question–answer and CoT data plus replay - Eight loops in Stage 1a caused "loss spikes and gradient oscillations", so they dropped to 4.
- Batches ramped from 4M to 8M tokens for stability.
- "Recurrent architectures require smaller learning rates than parameter-matched Transformers." No exhaustive sweeps were run.
- Reasoning SFT on 8.3M examples gave Ouro-Thinking. RL attempts (DAPO, GRPO) gave no gain.
-
Evaluations:
- base models against Qwen2.5/3, Gemma3 and Llama3.x bases, with lm-eval-harness and evalplus;
- Thinking models against Qwen3 and DeepSeek-Distill, with an in-house harness and an LLM judge;
- performance by loop count, T = 1–8;
- early-exit strategies and KV-cache sharing;
- safety and faithfulness;
- controlled synthetic "physics of LMs" experiments on knowledge capacity against knowledge manipulation;
- a small-scale scaling study (53M–1.36B parameters, 20B tokens).
Main results (verified)
The parameter-efficiency claim, in the authors' words. "1.4B and 2.6B parameter LoopLMs match 4B and 8B standard transformers on most benchmarks, yielding 2-3× parameter-efficiency gains." The abstract also says "match the results of up to 12B SOTA LLMs".
- Ouro-1.4B (4 loops, 7.7T tokens) against Qwen3-4B base (dense, 36T tokens):
- higher on BBH 71.02 against 70.95, GSM8K 78.92 against 72.86, MATH500 82.40 against 59.60, and Winogrande;
- lower on MMLU 67.35 against 73.19, MMLU-Pro 48.62 against 51.40, HumanEval 74.40 against 77.40, and MBPP, ARC-C and HellaSwag.
- Against the similar-size Qwen3-1.7B: MMLU 67.35 against 62.46, BBH 71.02 against 53.51, MATH500 82.40 against 25.80.
- Ouro-2.6B against Qwen3-8B base:
- higher on MMLU-Pro 55.73 against 53.72, BBH 80.46 against 77.65, MATH500 90.85 against 62.30;
- lower on MMLU 74.60 against 76.63, GSM8K 81.58 against 83.09, HumanEval 78.70 against 84.80.
- The gains are largest on reasoning-heavy benchmarks and smallest on knowledge-heavy ones. The authors say so, and the MMLU breakdown shows it (below).
- Thinking models (LLM judge): Ouro-1.4B-Thinking scored OlympiadBench 71.55 against Qwen3-4B's 73.18, and BeyondAIME 34.0 against 31.0. The 2.6B scored 76.44 against Qwen3-8B's 75.25, and 39.0 against 38.0.
Loop count.
- The base 1.4B's MMLU by loop: 41.21 at T = 1, then 60.43, 66.71 and 67.45 at T = 4. Beyond the trained depth it degrades: 66.64 down to 64.49 at T = 5–8.
- The Thinking models are near zero at T = 1 (AIME24 0.00) and peak at T = 4–5. AIME24 at T = 8 falls to 38.67, from 65.00.
- Performance is tied to the trained depth; extrapolating to more loops helps nothing.
Knowledge capacity against knowledge manipulation (section 6, Appendix B).
- Capacity.
- Setup: synthetic biographies bioS(N), N = 20K–500K people, about 47.6 bits each, 1,000 exposures. GPT-2-style models with RoPE, about 1M to 40M parameters, trained with 1 or 4 loops.
- "Looping does not increase knowledge capacity nor improve capacity scaling": both attain about 2 bits per parameter.
- Manipulation, Mano task.
- Setup: modular-arithmetic expression trees mod 23, answered with no intermediate tokens, at lengths L = 10, 16 and 24. Looped k⊗12/k models (k = 2, 3, 6) were compared with dense {2, 3, 6, 12}⊗1 models.
- Looped models "always outperform their non-looped counterpart" at equal parameters, and are "better or comparable" to the iso-FLOP 12-layer model.
- The figure text layer, mapping mine, gives at L = 24: 3 layers 11.0% against the same 3 layers looped 4 times 92.2%; 12 dense layers 34.8%.
- Manipulation, multi-hop QA.
- Setup: 3-hop questions over 500 synthetic people and 20 relations. A 6-layer model was looped 1, 2 or 4 times, against a 24-layer iso-FLOP model.
- "Models with more loops require fewer samples", for example with 15% of question–answer pairs (12,000).
- MMLU by category, loop 1 against loop 4 of the same model.
- Largest gains: elementary mathematics +155.6%, formal logic +143.3%, logical fallacies +127.8%.
- Smallest: global facts +8.3%, moral scenarios +7.8%, virology +13.7%.
- Theory. A one-layer transformer looped O(log D) times can decide reachability on a graph partly stored in its weights and partly given in context. Continuous CoT needs O(D) steps and discrete CoT O(n²). The construction needs hidden width Θ(n).
Adaptive exit and the cache.
- On MMLU, accuracy rose from about 40% at 1 loop to about 60% at 2 and 67.35% at 4.
- The trained gate reached 66% at an average of 2.5 loops; the untrained gate about 64%.
- Prefill needs every loop's KV cache; reusing it cost more than 10 points on GSM8K.
- When decoding, keeping only the last loop's cache cut cache memory 4× at GSM8K 78.85 against 78.92 (MATH-500 80.40 against 82.40). Keeping only the first loop's collapsed (18.73).
The small-scale study (Appendix D: 53M, 134M, 374M, 778M, 1.36B; 20B FineWeb-Edu tokens; T = 1, 2, 4, 8).
- At T = 1 the LoopLM and the standard model are identical. In general, more loops help both.
- Exception: the LoopLM at 778M and 1.36B, where more loops did not always help.
- "The performance of the standard model exceeds that of LoopLM under the same conditions" in every case. Here "standard" means an unshared model with T times the layers: the same compute, more parameters.
- The gap (Table 18, average benchmark score) is 0.021–0.023 at 2 steps and 0.025–0.039 at 4 steps.
- It grows with steps and shrinks with size.
- "When the model size is insufficient, the shallow step-wise loss increases with the growing amount of training data." Small models sacrifice the early loops. "A larger model size may be more effective for LoopLM."
- Loss follows a Chinchilla-style law in N, D and the maximum number of steps (R² 0.96).
Other claims.
- Harmfulness falls as loops increase.
- Intermediate loops' answers disagree with the final one: step 2 agrees with step 4 on only 36.1% of Quora pairs. The authors read this as faithful latent reasoning.
- They propose using early loops as built-in draft models.
Limits
- Confounded comparisons. Every external baseline differs in data, tokens (Ouro 7.7T against Qwen3 36T), tokenizer and recipe. Ouro's Stage 4 includes 300B tokens of question–answer and CoT data, which plausibly explains much of its MATH500 margin over base models (82.40 against 25.80 for Qwen3-1.7B). No same-data, same-parameter dense twin is trained at 7.7T.
- Compute is not matched at inference. Four loops make each token cost about 4× an equal-parameter dense model's forward FLOPs. That is more than the 4B model the 1.4B matches.
- Independent reproduction failed (MELT, Appendix E). The MELT authors and a second group could not reproduce the reported scores. In their setup Ouro-2.6B-Thinking did not beat Qwen3-4B. Ouro-1.4B still beat same-size dense models.
- The released code runs all loops even when the exit gate fires, so adaptive exit saved no compute in practice.
- Last-loop cache sharing gave 0.0 on long-form reasoning.
- The knowledge-capacity and manipulation experiments are small (about 1M–40M parameters and toy tasks). The Mano result reports the best of 3 seeds. The 12-layer iso-FLOP baseline was not always beaten.
- Internal inconsistencies:
- Table 18 lists sizes 170M, 340M, 680M and 1.3B, while the text says 53M, 134M, 374M, 778M and 1.36B.
- Appendix B.2 calls the expression lengths "model depths".
- Thinking-model scores come from an LLM judge in an in-house harness.
What it means for Kurisutina (inference)
- Looping adds manipulation, not storage.
- At fixed parameters it does not raise capacity (about 2 bits per parameter, as in
summaries/base/allenzhu2024_knowledge_capacity.md). - It improves composing stored facts (multi-hop), symbolic procedures, and the sample efficiency of learning them. That is the capacity the R4 synthesis §7 places in the base: depth of processing, not knowledge.
- At fixed parameters it does not raise capacity (about 2 bits per parameter, as in
- "2–3×" is the authors' own parameter-efficiency claim; the user's "4–5×" is not in the paper.
- It is measured against dense models trained on different data. Compute at inference is about T = 4 times.
- At equal compute, an unshared deeper model beat a LoopLM in every small-scale comparison the paper ran.
- What looping saves is parameters, and so weight memory and knowledge storage; it does not save FLOPs.
- For a person-scaled "deliberation steps" setting:
- Accuracy rises steeply from 1 to 4 loops and degrades beyond the trained depth. A loop count can lower capability below the trained maximum (1–3 loops), but cannot raise it above training.
- This matches "settings lower but do not add". Whether loop-limited errors look like a weaker human's errors is untested (see Maia,
mcilroyyoung2020_maia.md).
- The gains appeared at 7.7T tokens. At 53M–1.36B on 20B tokens, the paper's own study shows small or absent gains at larger sizes, and a deficit against iso-FLOP dense models. A looped arm at our scale (125M–350M on 2–8B tokens) is a test, not an expected win.
- Training cost.
- Training FLOPs scale with loops: Stage 1a ran 8 loops over 3T tokens.
- Training stability needed fewer loops, larger batches and lower learning rates.
- Activation memory in backpropagation scales with layers × loops unless recomputed. My inference; the paper reports no training memory.