Kurisutina

Memory-Efficient Looped Transformer (MELT): decoupling compute from memory in looped language models

Victor Conchello Vendrell, Arnau Padrés Masdemont (equal contribution), Niccolò Grillo, Jordi Ros-Giralt, Arash Behboodi, Fabio Valerio Massoli; Qualcomm AI Research. arXiv:2605.07721v3 [cs.CL], 28 September 2026 (preprint; the first version was May 2026). Read by researcher R6 (capability), 6 October 2026, at the user's request. Provenance: papers/capability/melt2026_looped_memory.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (1,917 lines, 22 pages). This covers the abstract, sections 1–6, references, and Appendices A–G:
    • A: related work;
    • B: untrained KV-sharing failures;
    • C: hyperparameters and compute;
    • D: inference and training efficiency;
    • E: reproducibility of Ouro, and early exit;
    • F: theory;
    • G: assets.
  • Not read: Figures 1, 2 and 4 as images (captions read). pdftotext scrambled the layout of Table 1. I reconstructed its columns by recomputing each model's reported average from its row values. Every average matched: Ouro math 62.3, MELT math 61.5, Ouro non-math 48.6, MELT non-math 51.6.
  • The version matters. v3 describes the conversion as chunk-wise training plus distillation at every loop. It describes no "interpolated transition" phase. Progressive growing, which interpolates between old and new components, appears only in related work, as "most closely related". Earlier versions were not read.

What they did

  • The problem. A looped LM (LoopLM, here Ouro) applies the same N-layer stack T times per token. A standard implementation keeps a KV cache per layer and per loop, so cache memory is O(N × L × T).
  • MELT's change.
    • Each layer keeps one latent state per token, updated across loops by an element-wise learned gate: h_t = z_t ⊙ h_{t−1} + (1 − z_t) ⊙ x_t, with z_t = σ(x_t W_z + h_{t−1} U_z + b_z).
    • Keys and values are projected from h. One cache row is added per token, at the first loop.
    • Only the last W generated tokens keep per-loop keys and values (a sliding window, W = 50).
    • Memory becomes O(N × L + W × L × T).
    • The gates add about 0.2B parameters (24 layers × 2048² × 2). That is why Ouro-1.4B becomes MELT-1.6B, and Ouro-2.6B becomes MELT-3.1B.
  • Training.
    • Why ordinary parallel training fails. The cache entry for token t+1 needs token t's completed loops. That creates a sequential dependency across tokens, which neither standard transformers nor Ouro have.
    • Chunk-wise training. Sequences are processed in chunks of 500 tokens: in parallel within a chunk, sequentially across chunks.
    • Distillation at every loop, with the original Ouro as teacher.
    • Initialisation: Ouro-1.4B-Thinking or Ouro-2.6B-Thinking weights, with random gates. At first this gives incoherent output.
    • Data: 50% AceReason-1.1-SFT and 50% OpenThoughts3, about 32K samples, 320M tokens, 1,000 steps, sequence length 10K.
    • Loop count: 4 recurrent steps.
  • Compute.
    • The main run took 720 H100-hours (90 h on 8×H100).
    • Ablations took 1,440, evaluation about 500, and preliminary experiments about 15,000.
    • About 17,000 GPU-hours in all.
  • Evaluation:
    • Benchmarks: six math (AIME24/25/26, AMC23, MATH-500, OlympiadBench) and four general (GPQA, HLE subset of 300, MMLU-Redux, HumanEval+). Temperature 1.0, top-p 0.7, 32k-token completions.
    • Comparison models: Qwen3-1.7B, Gemma4-E2B, Qwen3.5-2B, DeepSeek-R1-1.5B, and Ouro itself.

Main results (verified)

  • MELT keeps Ouro's performance.

    • MELT-1.6B against Ouro-1.4B-Thinking, as the MELT authors measured it: math average 61.5 against 62.3; non-math 51.6 against 48.6.
    • AIME24 pass@1 50.4 against 50.2. AMC23 77.2 against 81.2.
    • MELT-3.1B against Ouro-2.6B: math 70.4 against 71.9; non-math 71.9 against 71.4.
  • Both beat same-size dense models on these reasoning benchmarks, with math averages 61.5–62.3 against 56.9 (Qwen3-1.7B), 56.0 (Gemma4-E2B), 46.9 (DeepSeek-R1-1.5B) and 40.7 (Qwen3.5-2B).

    • Exceptions: Qwen3-1.7B and Gemma4-E2B beat MELT on AMC23, and Qwen3.5-2B beats it on GPQA (45.1 against 42.9).
  • Memory for one 32k-token generation (H100; Table 3):

    Model Peak (GB) Weights (GB) KV cache (GB) KV per token (MB) Max batch
    Ouro-1.4B-Thinking 26.58 2.87 23.71 0.759 3
    MELT-1.6B 9.23 3.27 5.96 0.191 11
    Qwen3-1.7B (dense, GQA) 6.74 3.44 3.30 0.106 21
    • MELT cuts the KV cache 3.98× against Ouro.
    • It is still about 1.8× Qwen3's cache, "about 2.5 GB more", because MELT (like Ouro) has no grouped-query attention.
    • Under capped memory, MELT holds most of the accuracy–memory Pareto frontier (Figure 1, read from caption and text).
  • Speed (Table 10, batch 1):

    • decoding: MELT 8.5 tokens/s, Ouro 13.5, Qwen3 41.6;
    • prefill: 2,254, 1,796 and 6,974 tokens/s;
    • time to first token: 0.089, 0.111 and 0.029 s.
    • The authors: MELT supports a larger batch "at the cost of slower batch-1 decoding".
  • Training throughput (Table 11): chunk-wise MELT training ran 2,168 tokens/s against 20,646 for Ouro distillation in the same setting. That is about a 10× wall-clock slowdown.

  • Ablations.

    • Gates: element-wise gating averaged 64.2. Using only the last loop's state ("Last", no extra parameters) averaged 63.3. A plain mean of loops averaged 55.6.
    • Removing distillation at all loops dropped AIME24 pass@1 to 45.8.
    • Replacing chunk-wise training with fully parallel training dropped it to 10.6 (MATH-500 55.0). That is worse than no training at all: Ouro weights with random gates scored 14.2 (MATH-500 58.0). Respecting the sequential dependency in training is "indispensable".
  • Untrained cache sharing fails on long reasoning. Reusing the first or last loop's cache without training, the methods of Geiping 2025 and of the Ouro paper, gave 0.0 on every benchmark for Ouro-1.4B-Thinking. Generations start coherently and then degenerate. This contradicts the Ouro paper's report of near-parity on GSM8K and MATH-500. The MELT authors attribute that to Ouro's few-shot, format-constrained evaluation.

  • Reproducibility of Ouro (Appendix E, the authors' own findings):

    • "Despite substantial effort, we were unable to fully reproduce the reported results of Ouro." A second group (Gao et al. 2026) found the same.
    • In their setup, Ouro-2.6B-Thinking does not beat Qwen3-4B, contrary to the Ouro paper.
    • Ouro-1.4B-Thinking still beats same-size dense baselines.
  • Ouro's early exit does not save compute in the released code. The default threshold effectively disables it. Even when it triggers, all loops run anyway: the gate picks which loop's logits are used.

    • The authors' explanation: later tokens need the final loop's cache entries.
    • MELT's single cache could allow real early exit (future work).
  • Theory (Appendix F): where the gate saturates near 1, the state's Jacobian approaches the identity, so gradients are preserved across loops. The authors say this does not establish stability of the full recurrence.

Limits (stated or evident)

  • Only post-training conversion was tried.
    • "The current experiments adapt pretrained Ouro checkpoints; incorporating the MELT mechanism throughout pretraining remains future work."
    • "Given the minimal cost of adapting … and the additional training overhead introduced by chunk-wise training, we propose MELT as a lightweight post-training adaptation of an existing looped model."
    • From-scratch pretraining is neither tested nor ruled out.
  • Loop count fixed at 4. No GQA. One family (Ouro), two sizes, math and code reasoning data only. Single runs, with ± values from resampled completions.
  • "Same memory" means KV-cache memory that does not grow with loop count. Weights grow by the gates (+0.2B). Compute per token is still about T times a dense model's, which shows in 8.5 against 41.6 tokens/s.
  • The comparison baselines are other groups' released models, trained on different data. Nothing isolates looping from data.

What it means for Kurisutina (inference)

  • MELT changes memory, not capability. Its reasoning scores track Ouro's within a few points either way. Whatever a looped model buys, MELT keeps at a dense model's KV footprint. It buys nothing new itself, apart from making real early exit possible in future.
  • Deriving it from a LoopLM is the only tested path. For us that means pretraining our own LoopLM in parallel, then converting it. That stays within "we build our own base".
    • Pretraining MELT from scratch would pay the roughly 10× training slowdown on every pretraining token (my inference from Table 11). It would also have to learn under random gates, which the ablation shows fully parallel training cannot handle.
    • It is untested. As a first experiment it is a research project, not a recipe.
  • For a local 12 GB GPU, MELT's gain is in inference KV memory on long generations: 23.7 → 6.0 GB at 32k tokens for a 1.4B model.
    • Training memory is not reported.
    • Training a looped model keeps activations for every loop in the backward pass, so activation memory scales with loop count unless recomputed. This is my inference; neither paper measures it here.
  • The reproducibility notes are a caution on Ouro's headline claims. Independent measurement puts looped 1.4B models above same-size dense models on reasoning, but not at the larger-dense level Ouro reports at 2.6B.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.