Kurisutina

State-offset tuning: state-based parameter-efficient fine-tuning for state space models

arXiv:2503.03499v2 (9 June 2025), Seoul National University, FuriosaAI, University of Wisconsin–Madison and Ajou University. The venue is not printed in the arXiv text; a web search result lists it as an ACL 2025 short paper (that page was not read). Read by researcher R4f for research batch R4 (recurrence for the base and the person slot), 5 October 2026. Provenance: papers/base/kang2025_state_offset.provenance.json.

Why this paper

It was swapped in for Beck et al. 2024 (xLSTM).

  • R4b's section 7 infers, without a source, that a person carried as a recurrent state is overwritten, and that "the protected form is a parameter".
  • This paper is the nearest test in the literature: it conditions a pretrained SSM through its initial state, and through a constant offset applied at every step.
  • xLSTM's matrix memory is another fixed-size state, already covered by the bound in jelassi2024_copying.

What was read

  • Read in full: every line of the pdftotext conversion (14 pages, 2,728 lines): abstract, Sections 1–7 (with limitations and risks), references, Appendices A–G and Tables 1–13.
  • Not read: Figures 1–3, which are block diagrams; their captions and the equations in the text were read.
  • Table cells were rebuilt from a column-wise extraction, and every number used below was checked against its row and column.

What they did

  • Setting: parameter-efficient fine-tuning, restricted to the SSM module: S6 in Mamba 130M, 1.4B and 2.8B; SSD in Mamba-2 130M and 1.3B.
    • Tasks: GLUE (7 tasks), Spider (text to SQL), SAMSum (dialogue summarisation) and DART (data to text).
    • One run per configuration. Learning rates were picked by a one-epoch grid search on 1–2k examples, and the best validation epoch was reported (early stopping).
  • The argument:
    • With initial state h_0, an SSM gives h_t = Σ Ā^{t−i}·B̄·x_i + Ā^t·h_0. Prefix-tuning (adding learned virtual tokens) changes only the initial state, so it can do no more than tuning h_0 directly (citing Galim et al. 2024).
    • "Initial State Tuning" learns h_0 per channel. In S6 its effect is multiplied by ∏_{i≤t} A_i, which "tends to decrease over time, leading to inconsistent effects". The authors relate this to SSMs struggling to recall early tokens.
  • Their method, State-offset Tuning: add a learned constant at every step.
    • Variant (h): ŷ_t = y_t + C_t·h′. Variant (y): a per-channel output bias y′.
    • They show it is equivalent to "iterative suffix-tuning", which keeps the virtual tokens at the last position at every step.
  • Baselines:
    • LoRA (rank 8 on all weight matrices in S6; rank 16 in SSD), BitFit, Additional-scan;
    • Prompt Tuning, Prefix-Tuning, Initial State Tuning, SDT;
    • full fine-tuning of the SSM module only, and of all parameters.

Main results (verified)

  • A per-step offset beats the initial state.

    Task (score) Offset (h) Initial state LoRA (S6 only) Full FT, S6 only Full FT, all
    Mamba 1.4B, Spider (execution accuracy) 57.4 51.8 56.3 56.7 66.2
    Mamba 2.8B, Spider 65.0 59.7 63.9 65.7 71.8
    Mamba-2 1.3B, Spider 58.5 (low-rank 60.5) 54.3 45.4 55.1 (SSD) 64.8
    Mamba 130M, GLUE average 78.5 77.4 78.3 79.3 80.5
    • Other baselines, Mamba 1.4B Spider: BitFit 51.3, Prompt Tuning 43.6, Prefix-Tuning 39.7. Mamba 130M GLUE: Prompt 63.8, Prefix 68.6.
    • The exception: on Mamba-2 1.3B SAMSum, the initial state beat the offset (ROUGE-1/2/L 50.4/26.4/42.3 against 48.8/24.7/40.5). Single GLUE tasks also go both ways (SST-2 92.4 against 91.9). The advantage is consistent on Spider, not on every task.
    • SAMSum and DART differences are small: SAMSum ROUGE-1 runs from 50.0 to 51.2 across the state-based methods, LoRA and full fine-tuning.
  • Prompt-style conditioning is weak on SSMs. Every state-based method beat every prompt-based one. On GLUE, Prefix-Tuning reached 85% of full fine-tuning for Mamba 130M and 89% for the Transformer Pythia 160M; the per-step offset reached 98%.

  • The offset is cheap.

    • It trains 0.01–0.45% of the parameters.
    • At 1.4B, training used 18.77 GB and 0.67 s per batch, against 22.99 GB and 0.80 s for LoRA.
    • Inference overhead is under 0.03% of FLOPs, against 0.4–0.9% for unmerged LoRA.
  • Compute: about 17,000 GPU-hours for the whole project (RTX 3090 below 1B parameters, H100 above).

Limits

  • It adapts the model to tasks, not to persons: no identity, no knowledge, text only.
  • Only the SSM module is tuned, and full-model fine-tuning is far ahead (Spider 66.2 against 57.4 at 1.4B). The LoRA baseline here covers only the SSM matrices, not the usual all-module LoRA.
  • The fading of the initial state's effect is argued from the equations. No measurement of effect against position in the sequence is reported.
  • Single runs with no variance reported; learning rates chosen on training loss; best validation epoch reported.

What it means for the person slot (inference)

  • A person loaded as an initial state fades. In an SSM, a learned prefix and an initial state are the same thing, and their influence is multiplied by every later step's decay. A person held this way is progressively overwritten by the conversation, which is the drift problem again.
  • The robust form is a per-step parameter. Conditioning applied at every step beat the initial state and about matched full SSM-module fine-tuning with far fewer parameters. This is the evidence for R4b's inference that a protected person state is a parameter.
  • An author or person tag given as a prefix (MeCo-style) would fade on a recurrent backbone. On such a backbone it should be injected at every step, or carried by the attention layers of a hybrid.
  • Tuning the SSM module alone has a ceiling. A person slot on a recurrent base would probably need projection adapters as well.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.