This file is the single source of truth for what MT-LNN has and has not demonstrated. Every row maps to a reproducible table in BENCHMARKS.md. If a marketing doc, slide, or README states a number that is not in the "Proven" section below (or is contradicted by the "Retracted / Null" section), that doc is wrong and this file wins.
Last reconciled: 2026-08-29.
MT-LNN is a streaming-state recurrent architecture whose one independently proven, hard-to-replicate result is that its fast-weight state supports cross-window / cross-session associative recall (0.56 mean accuracy) that attention and LoRA score exactly 0.000 on by construction.
| Claim | Number | Status | BENCHMARKS.md section |
|---|---|---|---|
| ❌ RETRACTED 2026-07-19 — both runs 2K-step (undertrained) vs a weak simple-ref baseline. Superseded by the convergence row below. | "20K convergence + modern baseline (2026-07-19)" | ||
| LM quality at convergence (20K steps, n=3, fp32, WikiText-103) | modern_transformer 78.86 ± 0.25 (144.1M) < mt_lnn 88.93 ± 0.33 (126.0M) < transformer 94.14 ± 0.78 (142.1M). All 9 runs stable. | "20K convergence + modern baseline (2026-07-19)" | |
| Modern-trunk recipes close the 11.3% gap (2026-09-06 confirmation, 20K steps, n=3, fp32, same P0 protocol) | base 89.62 (88.04/91.33/89.48) vs all_on 73.97 (73.33/74.76/73.83) — paired ΔPPL −14.7/−16.6/−15.6, 3/3 seeds, pre-registered verdict CONFIRM. all_on beats the modern Transformer baseline 78.86 by ~6.5%. The three missing pieces (SwiGLU gated-expansion FFN + QK-RMSNorm + depth-scaled residual init; ffn_swiglu/qk_norm/scaled_residual_init, all default-off bit-equivalent) were each individually positive at 2K screening and stack without conflict. |
✅ SUPPORTED — the gap was missing-recipe, not the liquid mixer itself. Honest caveat: all_on is 35% over-param vs the 144.1M anchor (195.1M); matched-param variant (ffn_expansion≈0.92 → ~148M) is the next check before any "default on" decision. | "Modern-trunk three missing pieces" in BENCHMARKS.md + docs/MODERN_TRUNK.md §3.5; JSON benchmarks/results/modern_trunk_screen_*_20000step*.json |
| 125M recurrent/liquid model trains stably at scale (the review's central "does it converge at 100x?" fear) | No NaN / no divergence for transformer / lnn / mt_lnn (fp16, 2026-07-05) and transformer / mt_lnn / mamba (fp32, 2026-07-16) at 125M | ✅ proven | "Scaling to ~125M → --mode train" + "fp32 reconfirmation + Mamba baseline (2026-07-16)" |
| First Mamba baseline at matched 125M-class param count | fp32, 2000 steps only, seed 0: MT-LNN 257.5 vs Mamba 414.0 PPL | "fp32 reconfirmation + Mamba baseline (2026-07-16)" | |
| Cross-window associative recall through fast-weight state | 0.56 mean (3-seed 0.621 / 0.434 / 0.621, ±0.09; in-window 0.99–1.00 every seed) | ✅ proven; attention/LoRA are 0.000 by construction (structural zero-channel) | "Cross-window associative recall" |
| The fast-weight matrix is the memory (not incidental) | Remove fast-weight → cross-window collapses 0.553 → 0.008; v1's 62.8M EMA state manages only 0.002 | ✅ proven | "Cross-window associative recall" |
| Cross-session snapshot → disk → fresh process → restore is lossless | Round-trip Δ +0.008 / +0.000; unit test bit-exact (max|diff| 0.0); wrong-session restore = chance; no-restore = chance | ✅ proven | "Cross-session persistence" |
| O(1) inference memory — O-series (ARR) only, audited vs compressed KV | ARR state flat 0.381 MB vs KV-cache O(T): 3.9x @512 → 252x @32k → 1008x @128k → 8063x @1M (fp16 GQA=1; KV would be 3072 MB). Vs the strongest compression (2-bit + GQA=8): 1070.9x @128k, 8567x @1M; vs 2-bit GQA=1: 1071x @1M. Every non-evicted crossover sits at T* ∈ [17, 976] ≪ 128k — ARR is smaller throughout the long-context regime. Verified across a 2048x context increase with the state unchanged to the decimal. | ✅ proven vs all non-evicted KV (byte-exact frontier ledger, docs/KV_FRONTIER.md); ❌ NOT vs sink(4)+window(512) eviction @2-bit GQA=1 (0.202 MB flat = 0.53x ARR, never wins) — see "What we do NOT claim". ARR figure is a measured snapshot-byte sum; KV figures are exact analytic. | "Scaling to ~125M → --mode decode" + "KV-compression frontier (2026-08-29)" |
| Adapter (M-series SFT) is capability-neutral — recall machinery costs no core ability | LAMBADA +0.7pt, ARC-easy +0.7/+1.2, HellaSwag −0.7/−0.9, PIQA 0.0/+0.4 (all ±1pt, noise) | ✅ proven (deployment-safety) | "Capability evals — v2s SFT is ability-neutral" |
| Bio prior is a good initialization, not the endpoint | Frozen-τ cross-window 0.285 vs trained 0.621 (bio init ≈46% of the effect, 285x chance; training doubles it) | ✅ proven | "Cross-window recall → Bio-prior (frozen-τ) ablation" |
| Cloud-inject prompt template lifts factual accuracy on a real model | +13.3% (25/30 → 29/30 on Qwen-1.5B) — and is identical with vs without the MT adapter | ✅ proven for the template; the adapter contributes 0 to this (see "what we do NOT claim") | "Real cloud-inject numbers on Qwen-2.5-1.5B" |
| Selective Copy (toy, ~200K params, matched budget, fair decode) | MT-LNN seq-exact 0.895 vs Transformer 0.676 (×1.32) / LNN 0.727; ratio widens to ×2.0 at T=229 | ✅ proven at toy scale only; LNN baseline close behind — most gain is the liquid component | "Selective Copy" / "Long-context sweep" + "What this shows" |
| RETRACTED as an architectural-edge claim (2026-08-29): under the canonical GRU-D baseline (Che et al. 2018, Δt-fed like every arch) at the nearest legal width d_model=78, the pre-registered absorption test FIRED — |t(mt_lnn,gru_d)|=1.75<2 AND gru_d degrades LESS (+12.8pp vs mt_lnn +20.8pp); the old lstm/gru margin also did not reproduce at d=78. See NULL rows below. | ❌ RETRACTED — single-width artifact + absorbed by GRU-D; raw d=65 JSON retained as history | "Irregular-sampling streaming edge — GRU-D + multi-task" | |
| Irregular-sampling accuracy on real air-quality forecasting (UCI Beijing, held-out Dingling station, 10 seeds, natural + swept missingness, 2026-08-29 hardened) | At 60% extra drop, mt_lnn PM2.5 24.98 µg/m³ and TEMP 4.37 °C, best of 5 archs; Welch vs gru_d/lstm/gru/transformer: PM2.5 +6.51/+7.72/+5.40/+5.29, TEMP +5.20/+2.42/+4.35/+3.97 — ALL significant at n=10, including the transformer. | ✅ proven on this station (n=10; upgraded from the 5-seed pilot, effect strengthened) — but see the station-generalisation null below | "Irregular-sampling streaming edge — air quality"; benchmarks/results/airquality_irregular_10seed.json |
| Air-quality advantage across stations (held-out Gucheng, 5 seeds, same protocol) | At 60% drop ALL six archs converge (PM2.5 29.1–29.9, every |t|<1; TEMP same). mt_lnn is numerically best at the REGULAR tier but its irregular-robustness margin does not transfer. | ❌ NULL — station-generalisation: the air WIN is single-station (Dingling), not a cross-station law | "Irregular-sampling streaming edge — air quality"; benchmarks/results/airquality_irregular_gucheng.json |
| CPU/ONNX deployment profile of the Δt-input liquid regressor | ONNX parity 5.96e-08 (gate ≤1e-5; historical 3.58e-07), variable-Δt runtime schedules worst 1.19e-07; 26.0 µs/step single-core (d=78, T=128); state 3,120 B flat | ✅ proven; honest negatives inside: dynamic int8 does NOT shrink the liquid graph at this scale (1,616 vs 1,480 KiB fp32; lstm shrinks 3.5×), and O(1) state is any-RNN (lstm 1,248 B / gru 624 B, smaller) | "Irregular-sampling streaming edge — deployment profile" |
| Parametric fast-weight memory passes the four-competency memory bench at v0 (TinyLlama-1.1B, 3000 steps, 3 seeds/config) | D1 in-window retrieval mt_v2 0.9896 vs frozen baseline 0.5352 (lora_only 0.0026 / mt_v2_delta 0.0013 — both tuned controls collapsed below the frozen model). D2 cross-window recall after KV-drop: mt_v2 0.0911 — the only nonzero (baseline/lora_only structural 0.000 by construction). D4 real snapshot→disk→subprocess→restore: mt_v2 0.0677 (no-restore control 0.0). D3 conflict-resolution fails for every config (≤0.0013; the delta-rule "corrects the binding" prediction NOT reproduced at this budget) | ✅ proven at v0 scale (3 seeds) for D1/D2/D4 as mechanism-vs-structural-zero; D3 negative; absolute D2/D4 values modest — mechanism demonstrated, magnitude not | "Parametric memory four-competency bench" + benchmarks/results/parametric_memory_summary.json |
| Constant streaming state on the same task | mt_lnn flat at 2.6 KB across a 256× stream-length increase; transformer KV reaches 34 MB at 32K = 13,107× larger | ✅ proven — but note O(1) state is a property of any RNN: lstm (1,040 B) and gru (520 B) are also flat and smaller | "A2 edge pilot → streaming memory" |
| Battery SoH accuracy, regular sampling | transformer 0.0972 / mt_lnn 0.0986 / lstm 0.1034 / gru 0.1048 RMSE (Ah), 10 seeds; baseline 0.1572 | "A2 edge pilot → accuracy" | |
| Real recurrence (pscan) does real work | pscan vs legacy broadcast: seq-exact 0.965 vs 0.883 (+8.2pp), tok-acc 0.983 vs 0.942 | ✅ proven | "Parallel scan ablation" |
| Event-stream state estimation (synthetic DVS-physics streams, held-out episodes, 10 seeds) | mt_lnn nRMSE 0.974 / 0.750 / 0.978 at event densities θ=0.15/0.35/0.8 — the only architecture to beat the constant predictor (1.0) at every density; lstm/gru/transformer sit at 0.81–1.30. Citable cells: vs lstm at θ=0.15 (paired sign p<0.05, n=10), vs transformer at θ=0.15 and θ=0.35. At the widest tested Δt span (3.24 decades) mt_lnn 0.774±0.013 vs lstm 0.817±0.035, 9/10 seed-wins, p=0.0215 | ✅ proven — citable cells only; several cells archive-only (bimodality flags), see BENCHMARKS tables | "Event-native sensing streams" |
| Next-event-channel prediction on the same streams — citable NEGATIVE | At the dense tier (θ=0.15) mt_lnn 0.227 vs lstm 0.329 / gru 0.328 (chance 0.167) — mt_lnn significantly worse than both RNNs (both gates pass with mt_lnn losing) | ✅ proven (as a negative) — continuous-time integration helps carry state, not categorical next-event prediction | "Event-native sensing streams → Next-channel prediction" |
| Line | Status | Where |
|---|---|---|
Latent-recursion adjudication (stack vs core iterations; Coconut arXiv:2412.06769 / Huginn arXiv:2502.05171 alignment; our unique cell = continuous-time loop body via liquid_step_ladder) |
Task 2 adjudication complete 2026-08-31 — NULL (budget_wall). 48/48 configs on A100-80GB (30k steps × 6 seeds × 2 modes × d{1,2,4,8}); all 8 tiers at chance (stack d1→d8 = 0.0668→0.0675, non-monotonic, gain 0.0007 « 2σ 0.0048; core d8 = 0.0676). Preregistered judge_decision, unmodified: h_supported=False, diagnosis=budget_wall — the loop never learned the task at 30k steps, depth effects undecidable at this budget; criteria untouched. Task 3 (anytime frontier vs CoT) still pending GPU (A100 probe ≈41 h sequential). Solid output unchanged: analytic compute ledger — 8 latent stack iterations = 4.8× the FLOPs of 8 CoT tokens at probe scale (latent edge is memory: 0 KV + 4.1 KB constant state vs 104.8 KB KV peak), so the "saves tokens" narrative does not hold on FLOPs |
"Latent recursion" in BENCHMARKS.md + docs/LATENT_RECURSION.md |
| Old claim | The real number | Status | BENCHMARKS.md section |
|---|---|---|---|
| "MT adapter drops PPL −28.5% / −27.7% / −34.4% at 0.117–0.196% trainable params" (TinyLlama / Qwen-1.5B / Qwen-3B) | The MT adapter was frozen by PEFT — only LoRA trained. Controlled ablation: lora_only 7.984 vs mt_lora 7.920 (v1) / mt_v2_lora 7.918 — MT adds ≈0 PPL (−0.064 for +62.8M params, within noise). The "0.1–0.2% trainable" figures ARE the LoRA-only param counts. | ❌ RETRACTED (2026-07-04) | Correction note (2026-07-04) + "Attribution results" table |
| "First end-to-end evidence the MT-LNN inductive bias transfers to a real pretrained LM" / "gain grows with base size (−28% → −34%)" | The trend measures plain LoRA fine-tuning; the MT adapter transferred nothing measurable on in-window PPL. | ❌ RETRACTED | Correction note (2026-07-04) |
| "MT-LNN has O(1) working memory / long-context compression" (as a property of the hybrid M-series) | O(1) holds only for the attention-free O-series. The hybrid still contains attention → KV cache still grows → not O(1). Its training memory is a NEGATIVE (uses MORE than a plain Transformer at every length, OOMs at 4096 too, ~1.6x slower). | ❌ RETRACTED for the hybrid (real for O-series only) | "Scaling to ~125M → --mode profile" (negative) + "--mode decode" (O-series positive) |
| "State compresses long context / out-of-window LM gains" | NULL ×2: standard chunked-streaming −0.006 / −0.000; TBPTT state-carry +0.004 (noise). Full attention gains 0.89–0.94 PPL from 512→2048 that the state does not capture. The state is an episodic key→value memory, not compressed distributed context. | ❌ NULL | "Out-of-window streaming" + "State-carry (TBPTT) training" |
| "Orch-OR collapse gate / Φ̂ integration / anesthesia validation is a working consciousness biomarker and a unique selling point" | Modules are INERT in the trained path. AVP FAILED; Φ̂ rises under anesthesia (sign inverted vs theory). Φ̂ moves +8.499 signed only as a "hooks fire" artifact on toy activations with no real-data baseline (Kraskov bias at N=148, Lord et al.). | ❌ INERT / net liability — inspiration only, never load-bearing | "Anesthesia Validation Protocol" + "What this does NOT show" |
| "Optional bio modules (predictive coding, GWTB, world model, rhythm, Hebbian) improve quality" | All 5 are PPL-neutral at 48M (within ±0.3–0.5 noise band); the full stack costs 5.6% throughput for nothing; predictive coding (the one ON by default) trends negative. Lean core trunk is best. | ❌ NULL (archived behind flags) | "O1 module switch-matrix" |
| "Adapter improves long-context / needle retrieval" | Within the 2048 window base and adapter both ~0.87–1.0 (parity, inconclusive); at 4096 both 0.000 (base RoPE limit, not adapter). Attribution now shows the adapter adds nothing anyway. | "Needle-in-a-haystack (CORRECTED)" | |
| "ARR attention-free student matches teacher" | PPL 25.4 = 2.15× teacher (11.8), still falling — converging with tokens, not at parity. ARR cross-window recall is negative at current budget (curriculum retry queued). | "Round 2/3 distillation" + "ARR-student recall — negative" | |
| "An intermediate attention-backfill ratio dominates both 0:1 and 3:1 on the quality-memory Pareto" (H002, pre-registered gate: 1:4 ≥15% PPL improvement vs same-batch ratio-0) | Confirmation run (2026-09-06, stabilized recipe --warmup_b 200 after the fix in iter/o-series-hybrid-ratio e265875; 4 ratios × 6 seeds, 21/24 valid, single A100-80GB bf16): ratio 0 mean 142.89; 1:8 88.92 (+37.8%); 1:4 67.37 (+52.9% — gate PASS); 1:2 263.52 (3/6 arms diverged, −84.4%). 1:4 beats the control on 6/6 paired seeds. Memory stays two-column honest: 1:4 keeps 6/22 attention layers — 192.5 MB @32K fp16 vs 0.65 MB O(1) for ratio 0 (full table in benchmarks/results/rebuilt/arr_ratio_pareto.md). |
✅ SUPPORTED (r*=1:4; promotable, 6 seeds) | "O-series hybrid-ratio sweep" in BENCHMARKS.md; evidence benchmarks/arr_ratio_out/ + benchmarks/results/rebuilt/ (evidence branch evidence/arr-confirm-6seed); kb H002 |
"The liquid advantage on event streams widens as the Δt distribution stretches across decades" (pre-registered in event_dt_span.py before running) |
Judgement fixed before the run: PROVEN iff mt_lnn beats the strongest discrete baseline at the widest span through publishable() and the gap is wider there than at the narrowest span. Outcome (state task, 10 seeds): (a) passed — at 3.24 decades mt_lnn 0.774±0.013 vs lstm 0.817±0.035, 9/10 seed-wins, p=0.0215, best at every span; (b) failed — the gap NARROWS 0.141 → 0.044 as every architecture improves with span (slopes all negative: mt_lnn −0.043 vs lstm −0.100). |
❌ NULL on the headline trend; the significant widest-span advantage is the only citable piece | "Event-native sensing streams → Δt-span pressure test" |
| "Multi-timescale τ ladder is what makes the liquid core span-robust" (τ-interaction arm) | mt_lnn beats single-τ mt_lnn_tau1 in 9/10 seeds at the widest span (mean 0.774 vs 0.803), but the tau1 per-seed distribution is flagged bimodal (gap_ratio 3.16) → publishable() fails → not citable. |
❌ NULL (archive-only) | "Event-native sensing streams → Δt-span pressure test" |
| "Liquid core transfers its event-stream edge to real event-camera data" (N-Caltech101, held-out classes) | 4 GB no-auth download parsed in memory; 10 train / 10 held-out classes, next-event-polarity, 10 seeds: mt_lnn 0.545 ± 0.031 vs gru 0.520 ± 0.027, sign test 5W/3L/2T p=0.727 — no significant difference. | ❌ NULL — no real-data advantage is claimed | "Event-native sensing streams → Real data: N-Caltech101" |
| "Irregular-sampling robustness is an architectural edge over discrete RNNs" (battery, re-run with GRU-D, d=78, 10 seeds) | Pre-registered absorption test FIRED: |t(mt_lnn,gru_d)| 1.75 < 2 and gru_d degradation +12.8pp < mt_lnn +20.8pp; vs gru also not significant (+1.86). The 2026-07-24 d=65 margin (mt_lnn +7.7% best-of-RNNs) did not reproduce at d=78 (lstm degrades less than mt_lnn there) — width-fragile and absorbed by the canonical baseline. | ❌ NULL / claim retracted (2026-08-29) | "Irregular-sampling streaming edge — GRU-D + multi-task" |
| "Continuous-time inductive bias shows up where it is attributable" (synthetic Van der Pol control, density × jitter grid, 5 seeds) | At ALL high-irregularity (cv=1) tiers, mt_lnn vs gru_d/lstm/gru nowhere significant (t ∈ [−1.5, +0.6]); several tiers have gru_d/lstm/gru numerically ahead. mt_lnn's clear synthetic edge is at the REGULAR dense tier (0.0093 vs lstm 0.0291) — a regular-sampling strength, not the claimed mechanism. | ❌ NULL — evidence against the mechanism story | "Irregular-sampling streaming edge — synthetic control"; benchmarks/results/synth_ct_control.json |
| "Irregular-sampling robustness generalises across task domains" (pre-registered ≥2-of-3 rule) | Outcome: battery null (GRU-D absorbs) / air WIN / synthetic not won = 1/3 — NEGATIVE as pre-registered. Only the air-quality domain (n=5, one held-out station) supports the claim. | ❌ NULL overall; single-domain positive remains | "Irregular-sampling streaming edge — verdict"; benchmarks/results/ADJUDICATION_LOG.md |
| "Wiring Δt into the decay (λ_t = exp(−Δt_t/τ)) revives the continuous-time mechanism" (mt_lnn_dt probe, synth domain, R1/R2 pre-registered, 5 seeds, 2026-08-29) | CLOSED on both rules. R1 (in-distribution cv=1 tiers): 0/3 wins — vs gru_d/lstm/gru/mt_lnn all |t|<2, at dt0.15 significantly WORSE than gru (t=−2.09). R2 (Δt-shift tier, train 0.05→test 0.4): mt_lnn_dt 0.3005 beats lstm (+4.41) and feature-wired mt_lnn (+8.85) but LOSES to gru_d 0.2485 (−1.74) and gru (−1.49). The learned GRU-D decay is more shift-robust than the structural exp(−Δt/τ). | ❌ NULL — mechanism story closed; the CT narrative is not told anywhere in this repo | "Irregular-sampling streaming edge — mechanism probe"; JSON benchmarks/results/synth_ct_control_dt.json (evidence PR #23) |
| "Folding a GRU-D-style learnable decay into the liquid bank wins back the Δt-shift tier" (mt_lnn_ad probe, λ = g·exp(−Δt/τ) + (1−g)·MLP, synth domain, R2'/R1' pre-registered, 5 seeds, 2026-08-29) | ARCHIVED FOR GOOD on both rules. R2' (Δt-shift tier): mt_lnn_ad 0.3004 is statistically indistinguishable from the fixed-decay probe mt_lnn_dt 0.3005 (t=+0.00) — the learned path added nothing — and still loses gru_d 0.2485 (−1.85) and gru 0.2722 (−1.78); beats only lstm (+4.79). R1' (in-distribution cv=1): 0/3 tiers, behind gru_d/lstm/gru everywhere (t ∈ [−2.07, −0.84]). GRU-D's shift robustness does not transfer by gluing a learned λ onto the multi-scale bank. | ❌ NULL — double negative with mt_lnn_dt; decay direction permanently closed, no third probe | "Irregular-sampling streaming edge — adaptive-decay probe"; benchmarks/results/synth_ct_control_ad.json |
- We do NOT claim the MT adapter beats LoRA on perplexity. On in-window LM PPL it adds ≈0 beyond LoRA. The old −28/−34% numbers are retracted.
- We do NOT claim the hybrid (M-series) flagship is O(1). It contains attention; its KV cache grows and its training memory is worse than a Transformer's. O(1) is an inference property of the O-series (ARR) only.
- We do NOT claim long-context language-modeling gains. Out-of-window LM is a double null. The state carries discrete addressable key→value bindings, not compressed distributed context.
- We do NOT claim a working consciousness / integrated-information / Orch-OR result. Those modules are inert in the trained path and the AVP Φ̂ sign is inverted. Microtubule/Orch-OR framing is inspiration only, not evidence.
- We do NOT claim the optional bio modules improve quality. They are PPL-neutral; shipped configs run the lean core.
- We do NOT claim a 125M SOTA result. The 125M win is vs this repo's simple-reference Transformer and (as of 2026-07-16) a width/depth-mismatched Mamba baseline, single seed each, undertrained — a real, consistent, budget-limited signal, not a converged or SOTA-competitive result.
- We do NOT claim the ARR state is the minimum possible streaming memory. The KV-compression frontier ledger (2026-08-29) finds exactly one configuration that undercuts it at every context length: sink(4)+window(512) eviction with 2-bit GQA=1 KV — 0.202 MB flat, 0.53x the ARR state — and one de facto tie: the same eviction at 4-bit GQA=1 (1.03x). Eviction buys those bytes by discarding tokens: attention is structurally 0.000 on cross-window recall where the fast-weight state scores 0.56, so the two accounts are reported side by side, never merged. Correct claim: smallest carried state without discarding context.
- We do NOT claim an O(1) advantage at short context. Below the crossover T*, the (compressed) KV cache is simply smaller than the 0.381 MB state: T* ranges from 17 tokens (fp16 GQA=8) through 119 (2-bit GQA=8, the strongest compression) to 976 (2-bit GQA=1). The O(1) claim is a long-context claim; all crossovers sit three orders of magnitude below the 128k regime where the 1008x/8063x numbers are quoted.
- fp16 robustness: root-caused and FIXED (2026-07-19). MT-LNN's loss previously went non-finite mid-training under fp16 AMP (step 629 on Colab T4; reproduced locally at step 875/896 — the step drifts with data order). Instrumented diagnosis ruled out gradient explosion (grad-norm ~1.1 at failure) and activation overflow (peak 67.4 vs fp16's 65504 ceiling). Actual cause:
global_coherence.pycomputed(Q @ K^T) / scale, doing thed_head=64accumulation before the down-scaling — with measured q/k projection peaks of 52.8/60.8 the intermediate product hits ~2e5 and overflows fp16 inside the matmul, after whichInf * 0against the causal mask in_gate_energyyields NaN thatsigmoid()spreads through the layer. Fix: pre-scale ((Q/scale) @ K^T, 4 sites), select-instead-of-multiply in the gate reduction, fp32 accumulation, and aclamp_min(1e-6)replacing a1e-9epsilon that underflows to exactly 0 in fp16. Verified: the identical 2000-step recipe that died at step 875 now completes in full,stable: true, val PPL 257.91 vs Infinity — and matches fp32's 257.48 on the same recipe to within 0.17% (inside the ±4.89 seed variance), i.e. fp16 is back at fp32 parity, not just non-crashing. Audit: every other attention surface usesF.scaled_dot_product_attention;global_coherence.pywas the only hand-rolled one. Tool:benchmarks/diagnose_fp16_divergence.py. Long-horizon (20K-step) fp16 confirmation is still outstanding. - We do NOT claim the cloud-inject +13.3% as MT-adapter value. It is 100% the
[Absorbed fact]prompt template; the adapter row is identical to baseline.
Two checkpoints, one codebase, deliberately different trade-offs — kept split so every claim stays attributable (docs/PRODUCT_LINES.md):
- M-series — hybrid (attention + liquid adapter). Cloud/GPU serving where full base quality matters. Its unique edge is cross-window / cross-session associative recall (0.56 vs 0.000 structural for attention/LoRA) at ~1% parameter overhead, plus lossless cross-session state snapshot/restore. It is not a perplexity win over LoRA and not O(1).
- O-series — pure recurrent (ARR, attention-free). Edge / CPU / low-power / unbounded-stream where KV-cache growth is disqualifying. Its unique edge is genuine O(1) inference memory (flat 0.381 MB, 1008x smaller than KV at 128k). It is a research preview at 2.15× teacher PPL — a token-budget gap, not a stability problem. Quality escape hatch (H002, 2026-09-06): backfilling 6 of 22 layers with attention (ratio 1:4) cuts PPL to 1.6× teacher in the same budget (mean 67.37 vs 142.89 control, +52.9%) at 192.5 MB @32K — the O(1) line and the quality line are now endpoints of a measured curve, not a binary choice.
Everything else — the from-scratch 125M sample-efficiency result and the stability-at-scale result — is a shared architectural signal that motivates both lines.