Kurisutina

Reasoning with latent thoughts: on the power of looped transformers

Nikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar, Sashank J. Reddi; Google Research and TTI Chicago. ICLR 2025; arXiv:2502.17416v1 (24 February 2025). Read by researcher R6 (capability), 6 October 2026, as the source on what looping buys and what it does not. Provenance: papers/capability/saunshi2025_looped_reasoning.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (2,420 lines, 27 pages): abstract, sections 1–7, references, Appendix A (setups; Tables 5–8 per task), and Appendix B (notation, all proofs).
  • Not read: Figures 2–7 as images. Captions and text were read; the isoplots and scaling curves are described in the text.
  • Table 3's columns were matched to its stated "% Gap" formula, which I recomputed. One cell does not match: math word problems for 12⊗2 computes to 292%, against the printed 282%, probably rounding of unrounded averages.

What they did

  • Notation. (k ⊗ L) is a k-layer block looped L times. It has the parameters of the "iso-param" (k ⊗ 1) model and the FLOPs and depth of the "iso-FLOP" (kL ⊗ 1) model, which has L× the parameters.
  • Synthetic reasoning (tiny models, width 128–256), trained on input→answer with no chain of thought:
    • adding n three-digit numbers;
    • p-hop induction (alphabet 4, length 256);
    • i-GSM: symbolic grade-school problems, a DAG of depth 4, arithmetic mod 7.
    • Each was run with 3 seeds, reporting the best.
  • Language modelling.
    • 250B Pile tokens, with the same tokens in the same order for every model. The baseline is 24 layers, width 2048 ("1B" in the text, "1.5B" in Appendix A.2).
    • Compared: (k⊗24/k) looped models against (k⊗1) and (24⊗1), for k = 4, 6, 8, 12, plus a "middle looping" variant.
    • Evaluation: perplexity and 19 few-shot tasks in four groups. Closed-book QA tests memorisation; open-book QA, math word problems and "reasoning primitives" (variable-assignment chains) test reasoning.
  • "% Gap" is the share of the iso-param→iso-FLOP gap that the looped model covers.
  • Also:
    • isoplots of downstream score against perplexity during training;
    • scaling with effective depth D for (4⊗D/4) against (D⊗1);
    • a regulariser pulling the blocks of an ordinary model towards each other (cosine similarity);
    • theory.

Main results (verified)

Synthetic tasks: depth matters, parameters barely.

Task Shallow model Same model looped Deep model
Addition, 32 operands 1 layer: 0.0% 1⊗12: 99.6% 12 layers: 100%
Addition, 32 operands 2 layers: 38.8% 2⊗6: 99.5%
p-hop, p = 32 1 layer: 49.0% 1⊗6: 99.5% 6 layers: 99.6%
i-GSM 1 layer: 24.5% 1⊗8: 73.2% 8 layers: 73.2%
i-GSM 2 layers: 54.0% 2⊗4: 73.6%

On i-GSM, 1⊗2 and 1⊗4 reach 52.3% and 69.9%.

Language modelling (Table 3; scores are group averages):

Model Params, FLOPs Perplexity Closed-book QA Open-book QA Math word Reasoning primitives
24⊗1 baseline 24×, 24× 7.40 11.2 33.9 29.3 47.5
12⊗1 12×, 12× 8.16 8.2 26.9 26.7 35.7
12⊗2 12×, 24× 7.90 9.3 30.8 34.3 51.2
4⊗1 4×, 4× 10.12 1.8 13.8 9.7 19.4
4⊗6 4×, 24× 8.79 6.7 26.2 24.8 56.9
  • Looping recovers reasoning, not memorisation.
    • % Gap for perplexity is 34–48%, and for closed-book QA 37–58%.
    • For open-book QA it is 56–72%; for math word problems 77–282%; for reasoning primitives 131–153%. Every looped model beats the 24-layer baseline on the primitives, with up to 6× fewer parameters.
  • Isoplots. At the same perplexity, closed-book QA is the same for looped and baseline models; open-book QA and math are higher for looped ones. "Log perplexity is a very strong indicator of downstream performance on memorization based tasks."
  • Scaling. Accuracy ≈ α log D + β for both looped and deep models. α_loop/α_base is higher for reasoning groups and 1.19× for the primitives.
  • Middle looping (4 independent, 4 looped 4×, 4 independent layers, iso-param with 12⊗1) gives perplexity 7.81, closed-book 11.0 and primitives 56.5. That is more uniform than plain looping.
  • The regulariser. Applied to a 24-layer model (k = 4, λ = 10), it left perplexity unchanged (7.38) and raised math word problems to 36.4 (from 29.3), the primitives to 57.2 (from 47.5), and the all-task average to 30.0 (from 26.0). Blocks ended at cosine similarity ≈ 0.98.
  • Theory.
    • A 1-layer block looped ⌈log₂ n⌉ times composes n group elements.
    • A looped 1-layer transformer can simulate any L-layer model with R distinct layers, with width growing with R.
    • p-hop needs ⌊log₂ p⌋ + 2 loops.
    • A looped model with m loops can simulate m steps of chain of thought (Theorem 5.4). "CoT … is essentially a looped model that generates 1 thought token in each iteration"; looped models produce "multiple latent thoughts" per iteration.

Limits

  • The synthetic tasks are narrow and algorithmic, with tiny models, and report the best of 3 seeds ("we care about expressivity"). Language modelling is a single scale, about 1–1.5B, on one corpus.
  • Tasks are few-shot and small. Absolute scores are low (closed-book QA about 11%), so differences of a few points matter.
  • No inference cost: a looped model has the FLOPs of the deep model. There is no study of memory or KV cache.
  • Inconsistencies: "1B" against "1.5B" for the same baseline, and one % Gap cell (292 against 282). A typo: "deeper but shallower models".

What it means for Kurisutina (inference)

  • The cleanest evidence that storage and reasoning separate by architectural resource.
    • Parameters set what is memorised: closed-book QA tracks perplexity, and looping covers only about half the gap.
    • Depth, which looping supplies without parameters, sets multi-step reasoning: covered fully or exceeded.
    • This is the "capacity in the base" claim made concrete: reasoning depth is a property of compute per token, not of stored content.
  • For a fact-light engine (R4a), this argues that a small-parameter, deep-compute engine is the right shape. Its parameters need hold little knowledge (language plus common sense); its effective depth carries reasoning. A looped or middle-looped arm is a cheap test of that at our scale.
  • For person settings: accuracy grows like log(effective depth), so a loop count is a natural continuous dial for "how deeply this person reasons". Whether lowering it reproduces a person's errors is untested here.
  • For the slot: a slot that adds memorised content needs parameters, whether adapter or store. A slot that only shapes reasoning could be small. Looping cannot add stored knowledge.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.