Skip to content

fix(ai): preserve verified context windows across providers - #999

Merged
Drakonis96 merged 4 commits into
Drakonis96:mainfrom
bloosqr:pr/provider-context-windows
Sep 30, 2026
Merged

Drakonis96 merged 4 commits into
Drakonis96:mainfrom
bloosqr:pr/provider-context-windows

Conversation

@bloosqr

@bloosqr bloosqr commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Deep Research rejected prompts against models whose real context windows were missing from /models, or disappeared from the budget cache after five minutes. Preserve the last successful catalogue for the session and replace its snapshot on refresh, including smaller limits, removals and missing metadata. Isolate custom endpoints, including a configuration change while a catalogue request is pending. If custom discovery fails, retain its known windows while still offering manually entered model IDs.

Extend the original DeepSeek/Claude fix to Gemini (inputTokenLimit), Copilot (SDK catalogue refreshes), Cerebras (optional public metadata for authenticated models only), OpenAI, Xiaomi MiMo and OpenCode Go. Documented fallbacks match exact provider/model IDs; Go uses its own gateway limits, including smaller input ceilings. Unknown Claude IDs and Codex subscription models retain the conservative unknown-model cap. Local runtime allocations and advertised route limits take precedence. The Server catalogue uses the same generated shared contracts.

Validation:

  • The new integration regression fails against the original PR when a 600,000-token advertised window drops to 200,000 after five minutes, and passes with this change.
  • A second regression reproduces a custom discovery failure that incorrectly replaced an 8,192-token window with the 32,768-token unknown-model cap; it passes after preserving the successful snapshot.
  • Production catalogue parsing, budget resolution and both completion guards run under Electron. HTTP SDK traffic is confined to loopback fixtures; Copilot is stubbed at its runtime boundary. Tests cover refreshed reductions, failures, invalid metadata, endpoint isolation, pre-dispatch overflow rejection and Copilot's internal catalogue refresh.
  • AI/provider/context/subscription/retrieval regression suites passed, as did typecheck and lint. Server catalogue tests are included in root CI discovery; the generated module is checked against its TypeScript source.
  • Model limits were checked against provider documentation and public catalogues on 2026-09-30; all 33 Go entries were compared with the current provider-scoped registry. No paid inference requests were made.

bloosqr and others added 3 commits September 29, 2026 17:00
DeepSeek's /models reports no context windows, and only `deepseek-flash` was
known to have 1M. `deepseek-v4-flash` (a legacy name DeepSeek still accepts for
the same Flash model) fell back to 32,768, and with the byte-level prompt bound
a normal synthesis request was refused as a context overflow before it was sent.
The documented windows now live in one table covering every accepted name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…an 200K

The Anthropic model list was read without its context size, so a research
request to Claude Opus fell back to the 32K default and was refused before
sending ('The model does not have enough context'). Use the list's
max_input_tokens when present, with the documented 200K floor for Claude models.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Drakonis96 Drakonis96 changed the title fix(ai): know the context windows of DeepSeek and Claude models fix(ai): preserve verified context windows across providers Sep 30, 2026
@Drakonis96
Drakonis96 merged commit 846b630 into Drakonis96:main Sep 30, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants