Skip to content

Latest commit

 

History

History
99 lines (77 loc) · 30.4 KB

File metadata and controls

99 lines (77 loc) · 30.4 KB

MT-LNN Results — Canonical Evidence Base

This file is the single source of truth for what MT-LNN has and has not demonstrated. Every row maps to a reproducible table in BENCHMARKS.md. If a marketing doc, slide, or README states a number that is not in the "Proven" section below (or is contradicted by the "Retracted / Null" section), that doc is wrong and this file wins.

Last reconciled: 2026-08-29.


The one-paragraph honest summary

MT-LNN is a streaming-state recurrent architecture whose one independently proven, hard-to-replicate result is that its fast-weight state supports cross-window / cross-session associative recall (0.56 mean accuracy) that attention and LoRA score exactly 0.000 on by construction. ⚠️ The former headline "−31% validation PPL vs a matched Transformer" is RETRACTED as of 2026-07-19: it was measured at 2,000 steps (undertrained) against a simple-reference baseline. At convergence (20,000 steps, 3 seeds) the gap shrinks to 5.5% (88.93 ± 0.33 vs 94.14 ± 0.78), and a modern Transformer baseline (RoPE/RMSNorm/SwiGLU) beats MT-LNN by 11.3% (78.86 ± 0.25). Language-modeling perplexity is not currently an MT-LNN advantage — see below. A separate attention-free variant (the O-series / ARR) has genuinely O(1) inference memory (flat 0.381 MB vs an O(T) KV-cache, up to 1008x smaller at 128k context — an advantage re-audited against 2-bit / GQA-8 / sink+window eviction in the KV-compression frontier: it survives every non-evicted configuration and concedes small-window eviction on bytes, with the retrieval cost of that concession reported alongside). MT-LNN is a streaming-state recurrent architecture whose one independently proven, hard-to-replicate result is that its fast-weight state supports cross-window / cross-session associative recall (0.56 mean accuracy) that attention and LoRA score exactly 0.000 on by construction. ⚠️ The former headline "−31% validation PPL vs a matched Transformer" is RETRACTED as of 2026-07-19: it was measured at 2,000 steps (undertrained) against a simple-reference baseline. At convergence (20,000 steps, 3 seeds) the gap shrinks to 5.5% (88.93 ± 0.33 vs 94.14 ± 0.78), and a modern Transformer baseline (RoPE/RMSNorm/SwiGLU) beats MT-LNN by 11.3% (78.86 ± 0.25). Language-modeling perplexity is not currently an MT-LNN advantage — see below. A separate attention-free variant (the O-series / ARR) has genuinely O(1) inference memory (flat 0.381 MB vs an O(T) KV-cache, up to 1008x smaller at 128k context — an advantage re-audited against 2-bit / GQA-8 / sink+window eviction in the KV-compression frontier: it survives every non-evicted configuration and concedes small-window eviction on bytes, with the retrieval cost of that concession reported alongside). ⚠️ The 2026-07-24 "irregular-sampling robustness vs discrete RNNs" edge is likewise RETRACTED as an architectural claim (2026-08-29): under the canonical GRU-D baseline it is absorbed (pre-registered test fired) and the original margin did not survive the forced d_model 65→78 width migration; what survives is a single-domain result — air-quality forecasting, held-out station, where mt_lnn beats gru_d/lstm/gru and the transformer with |t|≥2.5 at high missingness (n=5). The earlier headline "−28–34% PPL adapter" wins were retracted (they were plain LoRA — the MT adapter was frozen and adds ≈0 PPL), out-of-window language modeling is a null result, and the Orch-OR / Φ̂ / consciousness modules are inert in the trained path (with the AVP sign inverted vs theory). We do not sell those. The earlier headline "−28–34% PPL adapter" wins were retracted (they were plain LoRA — the MT adapter was frozen and adds ≈0 PPL), out-of-window language modeling is a null result, and the Orch-OR / Φ̂ / consciousness modules are inert in the trained path (with the AVP sign inverted vs theory). We do not sell those.


PROVEN — reproducible, and the only claims allowed in public docs

Claim Number Status BENCHMARKS.md section
From-scratch native 125M MT-LNN beats a matched Transformer on val PPL fp16 AMP: 299.5 vs 435.6 (−31%); fp32: 257.5 vs 370.8 (−30.6%) ❌ RETRACTED 2026-07-19 — both runs 2K-step (undertrained) vs a weak simple-ref baseline. Superseded by the convergence row below. "20K convergence + modern baseline (2026-07-19)"
LM quality at convergence (20K steps, n=3, fp32, WikiText-103) modern_transformer 78.86 ± 0.25 (144.1M) < mt_lnn 88.93 ± 0.33 (126.0M) < transformer 94.14 ± 0.78 (142.1M). All 9 runs stable. ⚠️ honest negative: mt_lnn beats the simple baseline by 5.5% but loses to a modern Transformer by 11.3%. PPL is not an MT-LNN advantage. "20K convergence + modern baseline (2026-07-19)"
Modern-trunk recipes close the 11.3% gap (2026-09-06 confirmation, 20K steps, n=3, fp32, same P0 protocol) base 89.62 (88.04/91.33/89.48) vs all_on 73.97 (73.33/74.76/73.83) — paired ΔPPL −14.7/−16.6/−15.6, 3/3 seeds, pre-registered verdict CONFIRM. all_on beats the modern Transformer baseline 78.86 by ~6.5%. The three missing pieces (SwiGLU gated-expansion FFN + QK-RMSNorm + depth-scaled residual init; ffn_swiglu/qk_norm/scaled_residual_init, all default-off bit-equivalent) were each individually positive at 2K screening and stack without conflict. ✅ SUPPORTED — the gap was missing-recipe, not the liquid mixer itself. Honest caveat: all_on is 35% over-param vs the 144.1M anchor (195.1M); matched-param variant (ffn_expansion≈0.92 → ~148M) is the next check before any "default on" decision. "Modern-trunk three missing pieces" in BENCHMARKS.md + docs/MODERN_TRUNK.md §3.5; JSON benchmarks/results/modern_trunk_screen_*_20000step*.json
125M recurrent/liquid model trains stably at scale (the review's central "does it converge at 100x?" fear) No NaN / no divergence for transformer / lnn / mt_lnn (fp16, 2026-07-05) and transformer / mt_lnn / mamba (fp32, 2026-07-16) at 125M ✅ proven "Scaling to ~125M → --mode train" + "fp32 reconfirmation + Mamba baseline (2026-07-16)"
First Mamba baseline at matched 125M-class param count fp32, 2000 steps only, seed 0: MT-LNN 257.5 vs Mamba 414.0 PPL ⚠️ not re-run at convergence — given that the Transformer comparison reversed at 20K steps, this 2K-step Mamba result cannot be assumed to hold either. Treat as provisional; width/depth-mismatched external reference. "fp32 reconfirmation + Mamba baseline (2026-07-16)"
Cross-window associative recall through fast-weight state 0.56 mean (3-seed 0.621 / 0.434 / 0.621, ±0.09; in-window 0.99–1.00 every seed) ✅ proven; attention/LoRA are 0.000 by construction (structural zero-channel) "Cross-window associative recall"
The fast-weight matrix is the memory (not incidental) Remove fast-weight → cross-window collapses 0.553 → 0.008; v1's 62.8M EMA state manages only 0.002 ✅ proven "Cross-window associative recall"
Cross-session snapshot → disk → fresh process → restore is lossless Round-trip Δ +0.008 / +0.000; unit test bit-exact (max|diff| 0.0); wrong-session restore = chance; no-restore = chance ✅ proven "Cross-session persistence"
O(1) inference memory — O-series (ARR) only, audited vs compressed KV ARR state flat 0.381 MB vs KV-cache O(T): 3.9x @512 → 252x @32k → 1008x @128k → 8063x @1M (fp16 GQA=1; KV would be 3072 MB). Vs the strongest compression (2-bit + GQA=8): 1070.9x @128k, 8567x @1M; vs 2-bit GQA=1: 1071x @1M. Every non-evicted crossover sits at T* ∈ [17, 976] ≪ 128k — ARR is smaller throughout the long-context regime. Verified across a 2048x context increase with the state unchanged to the decimal. ✅ proven vs all non-evicted KV (byte-exact frontier ledger, docs/KV_FRONTIER.md); ❌ NOT vs sink(4)+window(512) eviction @2-bit GQA=1 (0.202 MB flat = 0.53x ARR, never wins) — see "What we do NOT claim". ARR figure is a measured snapshot-byte sum; KV figures are exact analytic. "Scaling to ~125M → --mode decode" + "KV-compression frontier (2026-08-29)"
Adapter (M-series SFT) is capability-neutral — recall machinery costs no core ability LAMBADA +0.7pt, ARC-easy +0.7/+1.2, HellaSwag −0.7/−0.9, PIQA 0.0/+0.4 (all ±1pt, noise) ✅ proven (deployment-safety) "Capability evals — v2s SFT is ability-neutral"
Bio prior is a good initialization, not the endpoint Frozen-τ cross-window 0.285 vs trained 0.621 (bio init ≈46% of the effect, 285x chance; training doubles it) ✅ proven "Cross-window recall → Bio-prior (frozen-τ) ablation"
Cloud-inject prompt template lifts factual accuracy on a real model +13.3% (25/30 → 29/30 on Qwen-1.5B) — and is identical with vs without the MT adapter ✅ proven for the template; the adapter contributes 0 to this (see "what we do NOT claim") "Real cloud-inject numbers on Qwen-2.5-1.5B"
Selective Copy (toy, ~200K params, matched budget, fair decode) MT-LNN seq-exact 0.895 vs Transformer 0.676 (×1.32) / LNN 0.727; ratio widens to ×2.0 at T=229 ✅ proven at toy scale only; LNN baseline close behind — most gain is the liquid component "Selective Copy" / "Long-context sweep" + "What this shows"
Irregular-sampling robustness on a real BMS task — "the liquid core's one demonstrated architectural edge over discrete RNNs" (NASA battery, 80% drop, mt_lnn +7.7% vs lstm +31.1% / gru +32.8%, t=+2.40/+2.63, d_model=65, 2026-07-24) RETRACTED as an architectural-edge claim (2026-08-29): under the canonical GRU-D baseline (Che et al. 2018, Δt-fed like every arch) at the nearest legal width d_model=78, the pre-registered absorption test FIRED — |t(mt_lnn,gru_d)|=1.75<2 AND gru_d degrades LESS (+12.8pp vs mt_lnn +20.8pp); the old lstm/gru margin also did not reproduce at d=78. See NULL rows below. ❌ RETRACTED — single-width artifact + absorbed by GRU-D; raw d=65 JSON retained as history "Irregular-sampling streaming edge — GRU-D + multi-task"
Irregular-sampling accuracy on real air-quality forecasting (UCI Beijing, held-out Dingling station, 10 seeds, natural + swept missingness, 2026-08-29 hardened) At 60% extra drop, mt_lnn PM2.5 24.98 µg/m³ and TEMP 4.37 °C, best of 5 archs; Welch vs gru_d/lstm/gru/transformer: PM2.5 +6.51/+7.72/+5.40/+5.29, TEMP +5.20/+2.42/+4.35/+3.97 — ALL significant at n=10, including the transformer. ✅ proven on this station (n=10; upgraded from the 5-seed pilot, effect strengthened) — but see the station-generalisation null below "Irregular-sampling streaming edge — air quality"; benchmarks/results/airquality_irregular_10seed.json
Air-quality advantage across stations (held-out Gucheng, 5 seeds, same protocol) At 60% drop ALL six archs converge (PM2.5 29.1–29.9, every |t|<1; TEMP same). mt_lnn is numerically best at the REGULAR tier but its irregular-robustness margin does not transfer. ❌ NULL — station-generalisation: the air WIN is single-station (Dingling), not a cross-station law "Irregular-sampling streaming edge — air quality"; benchmarks/results/airquality_irregular_gucheng.json
CPU/ONNX deployment profile of the Δt-input liquid regressor ONNX parity 5.96e-08 (gate ≤1e-5; historical 3.58e-07), variable-Δt runtime schedules worst 1.19e-07; 26.0 µs/step single-core (d=78, T=128); state 3,120 B flat ✅ proven; honest negatives inside: dynamic int8 does NOT shrink the liquid graph at this scale (1,616 vs 1,480 KiB fp32; lstm shrinks 3.5×), and O(1) state is any-RNN (lstm 1,248 B / gru 624 B, smaller) "Irregular-sampling streaming edge — deployment profile"
Parametric fast-weight memory passes the four-competency memory bench at v0 (TinyLlama-1.1B, 3000 steps, 3 seeds/config) D1 in-window retrieval mt_v2 0.9896 vs frozen baseline 0.5352 (lora_only 0.0026 / mt_v2_delta 0.0013 — both tuned controls collapsed below the frozen model). D2 cross-window recall after KV-drop: mt_v2 0.0911 — the only nonzero (baseline/lora_only structural 0.000 by construction). D4 real snapshot→disk→subprocess→restore: mt_v2 0.0677 (no-restore control 0.0). D3 conflict-resolution fails for every config (≤0.0013; the delta-rule "corrects the binding" prediction NOT reproduced at this budget) ✅ proven at v0 scale (3 seeds) for D1/D2/D4 as mechanism-vs-structural-zero; D3 negative; absolute D2/D4 values modest — mechanism demonstrated, magnitude not "Parametric memory four-competency bench" + benchmarks/results/parametric_memory_summary.json
Constant streaming state on the same task mt_lnn flat at 2.6 KB across a 256× stream-length increase; transformer KV reaches 34 MB at 32K = 13,107× larger ✅ proven — but note O(1) state is a property of any RNN: lstm (1,040 B) and gru (520 B) are also flat and smaller "A2 edge pilot → streaming memory"
Battery SoH accuracy, regular sampling transformer 0.0972 / mt_lnn 0.0986 / lstm 0.1034 / gru 0.1048 RMSE (Ah), 10 seeds; baseline 0.1572 ⚠️ all four statistically tied (every pairwise |t| < 1). MT-LNN reaches parity with 2.6× fewer params than the transformer — that, not accuracy, is the claim "A2 edge pilot → accuracy"
Real recurrence (pscan) does real work pscan vs legacy broadcast: seq-exact 0.965 vs 0.883 (+8.2pp), tok-acc 0.983 vs 0.942 ✅ proven "Parallel scan ablation"
Event-stream state estimation (synthetic DVS-physics streams, held-out episodes, 10 seeds) mt_lnn nRMSE 0.974 / 0.750 / 0.978 at event densities θ=0.15/0.35/0.8 — the only architecture to beat the constant predictor (1.0) at every density; lstm/gru/transformer sit at 0.81–1.30. Citable cells: vs lstm at θ=0.15 (paired sign p<0.05, n=10), vs transformer at θ=0.15 and θ=0.35. At the widest tested Δt span (3.24 decades) mt_lnn 0.774±0.013 vs lstm 0.817±0.035, 9/10 seed-wins, p=0.0215 ✅ proven — citable cells only; several cells archive-only (bimodality flags), see BENCHMARKS tables "Event-native sensing streams"
Next-event-channel prediction on the same streams — citable NEGATIVE At the dense tier (θ=0.15) mt_lnn 0.227 vs lstm 0.329 / gru 0.328 (chance 0.167) — mt_lnn significantly worse than both RNNs (both gates pass with mt_lnn losing) ✅ proven (as a negative) — continuous-time integration helps carry state, not categorical next-event prediction "Event-native sensing streams → Next-channel prediction"

PREREGISTERED / PENDING — protocol locked, no claims yet

Line Status Where
Latent-recursion adjudication (stack vs core iterations; Coconut arXiv:2412.06769 / Huginn arXiv:2502.05171 alignment; our unique cell = continuous-time loop body via liquid_step_ladder) Task 2 adjudication complete 2026-08-31 — NULL (budget_wall). 48/48 configs on A100-80GB (30k steps × 6 seeds × 2 modes × d{1,2,4,8}); all 8 tiers at chance (stack d1→d8 = 0.0668→0.0675, non-monotonic, gain 0.0007 « 2σ 0.0048; core d8 = 0.0676). Preregistered judge_decision, unmodified: h_supported=False, diagnosis=budget_wall — the loop never learned the task at 30k steps, depth effects undecidable at this budget; criteria untouched. Task 3 (anytime frontier vs CoT) still pending GPU (A100 probe ≈41 h sequential). Solid output unchanged: analytic compute ledger — 8 latent stack iterations = 4.8× the FLOPs of 8 CoT tokens at probe scale (latent edge is memory: 0 KV + 4.1 KB constant state vs 104.8 KB KV peak), so the "saves tokens" narrative does not hold on FLOPs "Latent recursion" in BENCHMARKS.md + docs/LATENT_RECURSION.md

RETRACTED / NULL / INERT — must not appear as selling points anywhere

Old claim The real number Status BENCHMARKS.md section
"MT adapter drops PPL −28.5% / −27.7% / −34.4% at 0.117–0.196% trainable params" (TinyLlama / Qwen-1.5B / Qwen-3B) The MT adapter was frozen by PEFT — only LoRA trained. Controlled ablation: lora_only 7.984 vs mt_lora 7.920 (v1) / mt_v2_lora 7.918 — MT adds ≈0 PPL (−0.064 for +62.8M params, within noise). The "0.1–0.2% trainable" figures ARE the LoRA-only param counts. ❌ RETRACTED (2026-07-04) Correction note (2026-07-04) + "Attribution results" table
"First end-to-end evidence the MT-LNN inductive bias transfers to a real pretrained LM" / "gain grows with base size (−28% → −34%)" The trend measures plain LoRA fine-tuning; the MT adapter transferred nothing measurable on in-window PPL. ❌ RETRACTED Correction note (2026-07-04)
"MT-LNN has O(1) working memory / long-context compression" (as a property of the hybrid M-series) O(1) holds only for the attention-free O-series. The hybrid still contains attention → KV cache still grows → not O(1). Its training memory is a NEGATIVE (uses MORE than a plain Transformer at every length, OOMs at 4096 too, ~1.6x slower). ❌ RETRACTED for the hybrid (real for O-series only) "Scaling to ~125M → --mode profile" (negative) + "--mode decode" (O-series positive)
"State compresses long context / out-of-window LM gains" NULL ×2: standard chunked-streaming −0.006 / −0.000; TBPTT state-carry +0.004 (noise). Full attention gains 0.89–0.94 PPL from 512→2048 that the state does not capture. The state is an episodic key→value memory, not compressed distributed context. ❌ NULL "Out-of-window streaming" + "State-carry (TBPTT) training"
"Orch-OR collapse gate / Φ̂ integration / anesthesia validation is a working consciousness biomarker and a unique selling point" Modules are INERT in the trained path. AVP FAILED; Φ̂ rises under anesthesia (sign inverted vs theory). Φ̂ moves +8.499 signed only as a "hooks fire" artifact on toy activations with no real-data baseline (Kraskov bias at N=148, Lord et al.). ❌ INERT / net liability — inspiration only, never load-bearing "Anesthesia Validation Protocol" + "What this does NOT show"
"Optional bio modules (predictive coding, GWTB, world model, rhythm, Hebbian) improve quality" All 5 are PPL-neutral at 48M (within ±0.3–0.5 noise band); the full stack costs 5.6% throughput for nothing; predictive coding (the one ON by default) trends negative. Lean core trunk is best. ❌ NULL (archived behind flags) "O1 module switch-matrix"
"Adapter improves long-context / needle retrieval" Within the 2048 window base and adapter both ~0.87–1.0 (parity, inconclusive); at 4096 both 0.000 (base RoPE limit, not adapter). Attribution now shows the adapter adds nothing anyway. ⚠️ INCONCLUSIVE / parity "Needle-in-a-haystack (CORRECTED)"
"ARR attention-free student matches teacher" PPL 25.4 = 2.15× teacher (11.8), still falling — converging with tokens, not at parity. ARR cross-window recall is negative at current budget (curriculum retry queued). ⚠️ research preview, not parity "Round 2/3 distillation" + "ARR-student recall — negative"
"An intermediate attention-backfill ratio dominates both 0:1 and 3:1 on the quality-memory Pareto" (H002, pre-registered gate: 1:4 ≥15% PPL improvement vs same-batch ratio-0) Confirmation run (2026-09-06, stabilized recipe --warmup_b 200 after the fix in iter/o-series-hybrid-ratio e265875; 4 ratios × 6 seeds, 21/24 valid, single A100-80GB bf16): ratio 0 mean 142.89; 1:8 88.92 (+37.8%); 1:4 67.37 (+52.9% — gate PASS); 1:2 263.52 (3/6 arms diverged, −84.4%). 1:4 beats the control on 6/6 paired seeds. Memory stays two-column honest: 1:4 keeps 6/22 attention layers — 192.5 MB @32K fp16 vs 0.65 MB O(1) for ratio 0 (full table in benchmarks/results/rebuilt/arr_ratio_pareto.md). ⚠️ Seed bimodality: 1:4 good mode 18.9–20.5 vs bad mode 162.1–164.0 (control also bimodal 74.5–207.6); warmup improved the bad mode (8/31 no-warmup batch: 220–305) but did not eliminate it — means must be cited with spread, paired comparison unaffected. ✅ SUPPORTED (r*=1:4; promotable, 6 seeds) "O-series hybrid-ratio sweep" in BENCHMARKS.md; evidence benchmarks/arr_ratio_out/ + benchmarks/results/rebuilt/ (evidence branch evidence/arr-confirm-6seed); kb H002
"The liquid advantage on event streams widens as the Δt distribution stretches across decades" (pre-registered in event_dt_span.py before running) Judgement fixed before the run: PROVEN iff mt_lnn beats the strongest discrete baseline at the widest span through publishable() and the gap is wider there than at the narrowest span. Outcome (state task, 10 seeds): (a) passed — at 3.24 decades mt_lnn 0.774±0.013 vs lstm 0.817±0.035, 9/10 seed-wins, p=0.0215, best at every span; (b) failed — the gap NARROWS 0.141 → 0.044 as every architecture improves with span (slopes all negative: mt_lnn −0.043 vs lstm −0.100). ❌ NULL on the headline trend; the significant widest-span advantage is the only citable piece "Event-native sensing streams → Δt-span pressure test"
"Multi-timescale τ ladder is what makes the liquid core span-robust" (τ-interaction arm) mt_lnn beats single-τ mt_lnn_tau1 in 9/10 seeds at the widest span (mean 0.774 vs 0.803), but the tau1 per-seed distribution is flagged bimodal (gap_ratio 3.16) → publishable() fails → not citable. ❌ NULL (archive-only) "Event-native sensing streams → Δt-span pressure test"
"Liquid core transfers its event-stream edge to real event-camera data" (N-Caltech101, held-out classes) 4 GB no-auth download parsed in memory; 10 train / 10 held-out classes, next-event-polarity, 10 seeds: mt_lnn 0.545 ± 0.031 vs gru 0.520 ± 0.027, sign test 5W/3L/2T p=0.727 — no significant difference. ❌ NULL — no real-data advantage is claimed "Event-native sensing streams → Real data: N-Caltech101"
"Irregular-sampling robustness is an architectural edge over discrete RNNs" (battery, re-run with GRU-D, d=78, 10 seeds) Pre-registered absorption test FIRED: |t(mt_lnn,gru_d)| 1.75 < 2 and gru_d degradation +12.8pp < mt_lnn +20.8pp; vs gru also not significant (+1.86). The 2026-07-24 d=65 margin (mt_lnn +7.7% best-of-RNNs) did not reproduce at d=78 (lstm degrades less than mt_lnn there) — width-fragile and absorbed by the canonical baseline. ❌ NULL / claim retracted (2026-08-29) "Irregular-sampling streaming edge — GRU-D + multi-task"
"Continuous-time inductive bias shows up where it is attributable" (synthetic Van der Pol control, density × jitter grid, 5 seeds) At ALL high-irregularity (cv=1) tiers, mt_lnn vs gru_d/lstm/gru nowhere significant (t ∈ [−1.5, +0.6]); several tiers have gru_d/lstm/gru numerically ahead. mt_lnn's clear synthetic edge is at the REGULAR dense tier (0.0093 vs lstm 0.0291) — a regular-sampling strength, not the claimed mechanism. ❌ NULL — evidence against the mechanism story "Irregular-sampling streaming edge — synthetic control"; benchmarks/results/synth_ct_control.json
"Irregular-sampling robustness generalises across task domains" (pre-registered ≥2-of-3 rule) Outcome: battery null (GRU-D absorbs) / air WIN / synthetic not won = 1/3 — NEGATIVE as pre-registered. Only the air-quality domain (n=5, one held-out station) supports the claim. ❌ NULL overall; single-domain positive remains "Irregular-sampling streaming edge — verdict"; benchmarks/results/ADJUDICATION_LOG.md
"Wiring Δt into the decay (λ_t = exp(−Δt_t/τ)) revives the continuous-time mechanism" (mt_lnn_dt probe, synth domain, R1/R2 pre-registered, 5 seeds, 2026-08-29) CLOSED on both rules. R1 (in-distribution cv=1 tiers): 0/3 wins — vs gru_d/lstm/gru/mt_lnn all |t|<2, at dt0.15 significantly WORSE than gru (t=−2.09). R2 (Δt-shift tier, train 0.05→test 0.4): mt_lnn_dt 0.3005 beats lstm (+4.41) and feature-wired mt_lnn (+8.85) but LOSES to gru_d 0.2485 (−1.74) and gru (−1.49). The learned GRU-D decay is more shift-robust than the structural exp(−Δt/τ). ❌ NULL — mechanism story closed; the CT narrative is not told anywhere in this repo "Irregular-sampling streaming edge — mechanism probe"; JSON benchmarks/results/synth_ct_control_dt.json (evidence PR #23)
"Folding a GRU-D-style learnable decay into the liquid bank wins back the Δt-shift tier" (mt_lnn_ad probe, λ = g·exp(−Δt/τ) + (1−g)·MLP, synth domain, R2'/R1' pre-registered, 5 seeds, 2026-08-29) ARCHIVED FOR GOOD on both rules. R2' (Δt-shift tier): mt_lnn_ad 0.3004 is statistically indistinguishable from the fixed-decay probe mt_lnn_dt 0.3005 (t=+0.00) — the learned path added nothing — and still loses gru_d 0.2485 (−1.85) and gru 0.2722 (−1.78); beats only lstm (+4.79). R1' (in-distribution cv=1): 0/3 tiers, behind gru_d/lstm/gru everywhere (t ∈ [−2.07, −0.84]). GRU-D's shift robustness does not transfer by gluing a learned λ onto the multi-scale bank. ❌ NULL — double negative with mt_lnn_dt; decay direction permanently closed, no third probe "Irregular-sampling streaming edge — adaptive-decay probe"; benchmarks/results/synth_ct_control_ad.json

What we do NOT claim

  • We do NOT claim the MT adapter beats LoRA on perplexity. On in-window LM PPL it adds ≈0 beyond LoRA. The old −28/−34% numbers are retracted.
  • We do NOT claim the hybrid (M-series) flagship is O(1). It contains attention; its KV cache grows and its training memory is worse than a Transformer's. O(1) is an inference property of the O-series (ARR) only.
  • We do NOT claim long-context language-modeling gains. Out-of-window LM is a double null. The state carries discrete addressable key→value bindings, not compressed distributed context.
  • We do NOT claim a working consciousness / integrated-information / Orch-OR result. Those modules are inert in the trained path and the AVP Φ̂ sign is inverted. Microtubule/Orch-OR framing is inspiration only, not evidence.
  • We do NOT claim the optional bio modules improve quality. They are PPL-neutral; shipped configs run the lean core.
  • We do NOT claim a 125M SOTA result. The 125M win is vs this repo's simple-reference Transformer and (as of 2026-07-16) a width/depth-mismatched Mamba baseline, single seed each, undertrained — a real, consistent, budget-limited signal, not a converged or SOTA-competitive result.
  • We do NOT claim the ARR state is the minimum possible streaming memory. The KV-compression frontier ledger (2026-08-29) finds exactly one configuration that undercuts it at every context length: sink(4)+window(512) eviction with 2-bit GQA=1 KV — 0.202 MB flat, 0.53x the ARR state — and one de facto tie: the same eviction at 4-bit GQA=1 (1.03x). Eviction buys those bytes by discarding tokens: attention is structurally 0.000 on cross-window recall where the fast-weight state scores 0.56, so the two accounts are reported side by side, never merged. Correct claim: smallest carried state without discarding context.
  • We do NOT claim an O(1) advantage at short context. Below the crossover T*, the (compressed) KV cache is simply smaller than the 0.381 MB state: T* ranges from 17 tokens (fp16 GQA=8) through 119 (2-bit GQA=8, the strongest compression) to 976 (2-bit GQA=1). The O(1) claim is a long-context claim; all crossovers sit three orders of magnitude below the 128k regime where the 1008x/8063x numbers are quoted.
  • fp16 robustness: root-caused and FIXED (2026-07-19). MT-LNN's loss previously went non-finite mid-training under fp16 AMP (step 629 on Colab T4; reproduced locally at step 875/896 — the step drifts with data order). Instrumented diagnosis ruled out gradient explosion (grad-norm ~1.1 at failure) and activation overflow (peak 67.4 vs fp16's 65504 ceiling). Actual cause: global_coherence.py computed (Q @ K^T) / scale, doing the d_head=64 accumulation before the down-scaling — with measured q/k projection peaks of 52.8/60.8 the intermediate product hits ~2e5 and overflows fp16 inside the matmul, after which Inf * 0 against the causal mask in _gate_energy yields NaN that sigmoid() spreads through the layer. Fix: pre-scale ((Q/scale) @ K^T, 4 sites), select-instead-of-multiply in the gate reduction, fp32 accumulation, and a clamp_min(1e-6) replacing a 1e-9 epsilon that underflows to exactly 0 in fp16. Verified: the identical 2000-step recipe that died at step 875 now completes in full, stable: true, val PPL 257.91 vs Infinity — and matches fp32's 257.48 on the same recipe to within 0.17% (inside the ±4.89 seed variance), i.e. fp16 is back at fp32 parity, not just non-crashing. Audit: every other attention surface uses F.scaled_dot_product_attention; global_coherence.py was the only hand-rolled one. Tool: benchmarks/diagnose_fp16_divergence.py. Long-horizon (20K-step) fp16 confirmation is still outstanding.
  • We do NOT claim the cloud-inject +13.3% as MT-adapter value. It is 100% the [Absorbed fact] prompt template; the adapter row is identical to baseline.

Product positioning (M-series vs O-series)

Two checkpoints, one codebase, deliberately different trade-offs — kept split so every claim stays attributable (docs/PRODUCT_LINES.md):

  • M-series — hybrid (attention + liquid adapter). Cloud/GPU serving where full base quality matters. Its unique edge is cross-window / cross-session associative recall (0.56 vs 0.000 structural for attention/LoRA) at ~1% parameter overhead, plus lossless cross-session state snapshot/restore. It is not a perplexity win over LoRA and not O(1).
  • O-series — pure recurrent (ARR, attention-free). Edge / CPU / low-power / unbounded-stream where KV-cache growth is disqualifying. Its unique edge is genuine O(1) inference memory (flat 0.381 MB, 1008x smaller than KV at 128k). It is a research preview at 2.15× teacher PPL — a token-budget gap, not a stability problem. Quality escape hatch (H002, 2026-09-06): backfilling 6 of 22 layers with attention (ratio 1:4) cuts PPL to 1.6× teacher in the same budget (mean 67.37 vs 142.89 control, +52.9%) at 192.5 MB @32K — the O(1) line and the quality line are now endpoints of a measured curve, not a binary choice.

Everything else — the from-scratch 125M sample-efficiency result and the stability-at-scale result — is a shared architectural signal that motivates both lines.