Reasoning effort controls - #6269
Conversation
|
Tested every provider end to end from the chat UI. For each one: set the effort with the brain icon in the chat input, sent a message, and checked the header label, the metrics line under the reply, and the
In every case the header, metrics line, and DB row all show the same effort that was sent. Provider-side confirmation that the levels are honored (direct SDK calls, same prompt):
Also verified:
|
|
fixes applied below: 1. Model-router workspaces couldn't undo a stored effort 2. An offline upgrade wiped the pricing cache 3. Ollama levels come from Ollama's own API 4. Sources
Both fail safely, since "Provider default" clears the choice. 5. Smaller fixes
|
- Add a reasoning effort picker to the chat prompt input, scoped per chat session (thread or workspace default chat) and kept in the browser only. - Validate the session's effort against the model's live capabilities before every request so unsupported levels are never sent. - Read cloud reasoning levels from models.dev and local provider levels from each provider's model API. - Map efforts to each provider's request format for chat and agent providers, and apply effort changes mid agent session. - Show OpenAI reasoning summaries and Gemini thoughts as thinking.
…line - Never apply a session reasoning effort on model-router workspaces. The picker is hidden there, so a stored effort could not be seen or cleared. - Keep the pricing cache when the reasoning cache is missing, so an offline upgrade does not lose cost data. - Read Ollama's per-model think values from /api/show, falling back to the model family on older Ollama versions. - Keep the home page's picker choice under its own draft session instead of the workspace's default chat. - Link each provider's reasoning wire format to its documentation and fix stale JSDoc.
ba3a60a to
b905294
Compare
- Read OpenRouter reasoning levels from models.dev: effort values, with "none" as off, and on/off for toggle models. Models with only a token budget, or no options (eg: mandatory thinking), get no controls. - Send the effort with OpenRouter's unified `reasoning` param on chat and agent requests, keeping the legacy `include_reasoning` flag when no effort is set.
Pull Request Type
Relevant Issues
None
Description
Adds a reasoning effort picker (brain icon) to the chat prompt input.
Scope: per chat session, per browser
Only verified levels are sent
Before every request, the stored level is checked against the model's current capabilities. A level the model does not list is dropped, so a value left over from a model switch can't fail the chat. A failed capability lookup also sends nothing, and is not cached so the next message retries. The picker only renders when the model lists levels.
Model router
Workspaces on the model router show no picker and never apply a stored effort, because the routed model changes per message.
Where each provider's levels come from
capabilities.effortoutput_config.effortreasoning_optionsreasoning.effort,reasoning.summary(summary dropped automatically for unverified orgs)reasoning_optionsthinking_config: a level for 3.x models, Google's documented budgets for 2.5reasoning_optionsthinking.type,reasoning_effort/api/showthinking.values, falling back to the model family on older Ollama versionsthink/api/v1/modelsreasoning.allowed_optionsreasoning_effort("on"is sent asmedium)reasoningmodel label, levels by model familychat_template_kwargsreasoning_options(effort levels and on/off; models with only a token budget, or mandatory thinking with no levels, get no picker)reasoning.effort/reasoning.enabled, replacing the legacyinclude_reasoningflag when an effort is setReasoning models that support it now show their reasoning (OpenAI summaries, Gemini thoughts) as a thinking block.
Paths
The effort applies to workspace and thread chats, including agent sessions, and can change mid agent session. The embed, developer API, OpenAI-compatible, and Telegram paths send no reasoning params, so the provider default applies.
Pricing cache
models.dev reasoning options are cached next to the existing pricing cache. A pricing cache from before this change keeps working offline, and the reasoning options are filled in on the next successful refresh.
Developer Validations
yarn lintfrom the root of the repo & committed changes