Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, Sashank J. Reddi; Google Research and TTI Chicago. ICLR 2025; arXiv:2502.17416v1 (24 February 2025). Read by researcher R6 (capability), 6 October 2026, as the source on what looping buys and what it does not. Provenance: papers/capability/saunshi2025_looped_reasoning.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (2,420 lines, 27 pages): abstract, sections 1–7, references, Appendix A (setups; Tables 5–8 per task), and Appendix B (notation, all proofs).
- Not read: Figures 2–7 as images. Captions and text were read; the isoplots and scaling curves are described in the text.
- Table 3's columns were matched to its stated "% Gap" formula, which I recomputed. One cell does not match: math word problems for 12⊗2 computes to 292%, against the printed 282%, probably rounding of unrounded averages.
What they did
- Notation. (k ⊗ L) is a k-layer block looped L times. It has the parameters of the "iso-param" (k ⊗ 1) model and the FLOPs and depth of the "iso-FLOP" (kL ⊗ 1) model, which has L× the parameters.
- Synthetic reasoning (tiny models, width 128–256), trained on input→answer with no chain of thought:
- adding n three-digit numbers;
- p-hop induction (alphabet 4, length 256);
- i-GSM: symbolic grade-school problems, a DAG of depth 4, arithmetic mod 7.
- Each was run with 3 seeds, reporting the best.
- Language modelling.
- 250B Pile tokens, with the same tokens in the same order for every model. The baseline is 24 layers, width 2048 ("1B" in the text, "1.5B" in Appendix A.2).
- Compared: (k⊗24/k) looped models against (k⊗1) and (24⊗1), for k = 4, 6, 8, 12, plus a "middle looping" variant.
- Evaluation: perplexity and 19 few-shot tasks in four groups. Closed-book QA tests memorisation; open-book QA, math word problems and "reasoning primitives" (variable-assignment chains) test reasoning.
- "% Gap" is the share of the iso-param→iso-FLOP gap that the looped model covers.
- Also:
- isoplots of downstream score against perplexity during training;
- scaling with effective depth D for (4⊗D/4) against (D⊗1);
- a regulariser pulling the blocks of an ordinary model towards each other (cosine similarity);
- theory.
Main results (verified)
Synthetic tasks: depth matters, parameters barely.
| Task | Shallow model | Same model looped | Deep model |
|---|---|---|---|
| Addition, 32 operands | 1 layer: 0.0% | 1⊗12: 99.6% | 12 layers: 100% |
| Addition, 32 operands | 2 layers: 38.8% | 2⊗6: 99.5% | |
| p-hop, p = 32 | 1 layer: 49.0% | 1⊗6: 99.5% | 6 layers: 99.6% |
| i-GSM | 1 layer: 24.5% | 1⊗8: 73.2% | 8 layers: 73.2% |
| i-GSM | 2 layers: 54.0% | 2⊗4: 73.6% |
On i-GSM, 1⊗2 and 1⊗4 reach 52.3% and 69.9%.
Language modelling (Table 3; scores are group averages):
| Model | Params, FLOPs | Perplexity | Closed-book QA | Open-book QA | Math word | Reasoning primitives |
|---|---|---|---|---|---|---|
| 24⊗1 baseline | 24×, 24× | 7.40 | 11.2 | 33.9 | 29.3 | 47.5 |
| 12⊗1 | 12×, 12× | 8.16 | 8.2 | 26.9 | 26.7 | 35.7 |
| 12⊗2 | 12×, 24× | 7.90 | 9.3 | 30.8 | 34.3 | 51.2 |
| 4⊗1 | 4×, 4× | 10.12 | 1.8 | 13.8 | 9.7 | 19.4 |
| 4⊗6 | 4×, 24× | 8.79 | 6.7 | 26.2 | 24.8 | 56.9 |
- Looping recovers reasoning, not memorisation.
- % Gap for perplexity is 34–48%, and for closed-book QA 37–58%.
- For open-book QA it is 56–72%; for math word problems 77–282%; for reasoning primitives 131–153%. Every looped model beats the 24-layer baseline on the primitives, with up to 6× fewer parameters.
- Isoplots. At the same perplexity, closed-book QA is the same for looped and baseline models; open-book QA and math are higher for looped ones. "Log perplexity is a very strong indicator of downstream performance on memorization based tasks."
- Scaling. Accuracy ≈ α log D + β for both looped and deep models. α_loop/α_base is higher for reasoning groups and 1.19× for the primitives.
- Middle looping (4 independent, 4 looped 4×, 4 independent layers, iso-param with 12⊗1) gives perplexity 7.81, closed-book 11.0 and primitives 56.5. That is more uniform than plain looping.
- The regulariser. Applied to a 24-layer model (k = 4, λ = 10), it left perplexity unchanged (7.38) and raised math word problems to 36.4 (from 29.3), the primitives to 57.2 (from 47.5), and the all-task average to 30.0 (from 26.0). Blocks ended at cosine similarity ≈ 0.98.
- Theory.
- A 1-layer block looped ⌈log₂ n⌉ times composes n group elements.
- A looped 1-layer transformer can simulate any L-layer model with R distinct layers, with width growing with R.
- p-hop needs ⌊log₂ p⌋ + 2 loops.
- A looped model with m loops can simulate m steps of chain of thought (Theorem 5.4). "CoT … is essentially a looped model that generates 1 thought token in each iteration"; looped models produce "multiple latent thoughts" per iteration.
Limits
- The synthetic tasks are narrow and algorithmic, with tiny models, and report the best of 3 seeds ("we care about expressivity"). Language modelling is a single scale, about 1–1.5B, on one corpus.
- Tasks are few-shot and small. Absolute scores are low (closed-book QA about 11%), so differences of a few points matter.
- No inference cost: a looped model has the FLOPs of the deep model. There is no study of memory or KV cache.
- Inconsistencies: "1B" against "1.5B" for the same baseline, and one % Gap cell (292 against 282). A typo: "deeper but shallower models".
What it means for Kurisutina (inference)
- The cleanest evidence that storage and reasoning separate by architectural resource.
- Parameters set what is memorised: closed-book QA tracks perplexity, and looping covers only about half the gap.
- Depth, which looping supplies without parameters, sets multi-step reasoning: covered fully or exceeded.
- This is the "capacity in the base" claim made concrete: reasoning depth is a property of compute per token, not of stored content.
- For a fact-light engine (R4a), this argues that a small-parameter, deep-compute engine is the right shape. Its parameters need hold little knowledge (language plus common sense); its effective depth carries reasoning. A looped or middle-looped arm is a cheap test of that at our scale.
- For person settings: accuracy grows like log(effective depth), so a loop count is a natural continuous dial for "how deeply this person reasons". Whether lowering it reproduces a person's errors is untested here.
- For the slot: a slot that adds memorised content needs parameters, whether adapter or store. A slot that only shapes reasoning could be small. Looping cannot add stored knowledge.