Skip to content

Reasoning effort controls - #6269

Merged
timothycarambat merged 4 commits into
masterfrom
feat/reasoning-controls
Oct 1, 2026
Merged

timothycarambat merged 4 commits into
masterfrom
feat/reasoning-controls

Conversation

@shatfield4

@shatfield4 shatfield4 commented Sep 2, 2026 •

Copy link
Copy Markdown
Collaborator

Pull Request Type

  • ✨ feat (New feature)
  • 🐛 fix (Bug fix)
  • ♻️ refactor (Code refactoring without changing behavior)
  • 💄 style (UI style changes)
  • 🔨 chore (Build, CI, maintenance)
  • 📝 docs (Documentation updates)

Relevant Issues

None

Description

Adds a reasoning effort picker (brain icon) to the chat prompt input.

Scope: per chat session, per browser

  • The choice is stored in the browser's local storage, keyed by workspace and thread (or the workspace's default chat). It is sent with each chat request and is never saved to the workspace, the thread, or the chat history. In multi-user mode, one user's choice never changes another user's chats, and a choice on one thread does not carry to other threads.
  • The home page's picker keeps its choice under its own draft session until the first message creates a thread. The choice then moves to that thread.
  • "Provider default" clears the choice and sends no reasoning params.
  • There is no workspace-level or system-wide reasoning setting.

Only verified levels are sent
Before every request, the stored level is checked against the model's current capabilities. A level the model does not list is dropped, so a value left over from a model switch can't fail the chat. A failed capability lookup also sends nothing, and is not cached so the next message retries. The picker only renders when the model lists levels.

Model router
Workspaces on the model router show no picker and never apply a stored effort, because the routed model changes per message.

Where each provider's levels come from

Provider Supported levels Request format
Anthropic Models API capabilities.effort output_config.effort
OpenAI models.dev reasoning_options reasoning.effort, reasoning.summary (summary dropped automatically for unverified orgs)
Gemini models.dev reasoning_options thinking_config: a level for 3.x models, Google's documented budgets for 2.5
DeepSeek models.dev reasoning_options thinking.type, reasoning_effort
Ollama /api/show thinking.values, falling back to the model family on older Ollama versions think
LM Studio /api/v1/models reasoning.allowed_options reasoning_effort ("on" is sent as medium)
Lemonade reasoning model label, levels by model family llama.cpp chat_template_kwargs
OpenRouter models.dev reasoning_options (effort levels and on/off; models with only a token budget, or mandatory thinking with no levels, get no picker) reasoning.effort / reasoning.enabled, replacing the legacy include_reasoning flag when an effort is set

Reasoning models that support it now show their reasoning (OpenAI summaries, Gemini thoughts) as a thinking block.

Paths
The effort applies to workspace and thread chats, including agent sessions, and can change mid agent session. The embed, developer API, OpenAI-compatible, and Telegram paths send no reasoning params, so the provider default applies.

Pricing cache
models.dev reasoning options are cached next to the existing pricing cache. A pricing cache from before this change keeps working offline, and the reasoning options are filled in on the next successful refresh.

Developer Validations

  • I ran yarn lint from the root of the repo & committed changes
  • Relevant documentation has been updated (if applicable)
  • I have tested my code functionality
  • Docker build succeeds locally

@shatfield4 shatfield4 self-assigned this Sep 2, 2026
@shatfield4

Copy link
Copy Markdown
Collaborator Author

Tested every provider end to end from the chat UI. For each one: set the effort with the brain icon in the chat input, sent a message, and checked the header label, the metrics line under the reply, and the workspace_chats row. To confirm the effort actually reached the provider, I ran the dev server with a temporary fetch tap that logged the reasoning fields of every outbound request body (not committed, only used for testing).

Provider / model Effort set in UI Request body sent to provider Reply
OpenAI gpt-5.1 high / off reasoning: { effort: "high" } / { effort: "none" } 42 vs 13 output tokens, correct
Anthropic claude-sonnet-5 max output_config: { effort: "max" } correct
Gemini gemini-3.5-flash minimal reasoning_effort: "minimal" correct
DeepSeek deepseek-flash on thinking: { type: "enabled" } thoughts block shown, correct
Ollama qwen3:8b off think: false no thinking, correct
LM Studio grayline-qwen3-8b off reasoning_effort: "none" no thinking, correct
Lemonade Qwen3-4B-GGUF on chat_template_kwargs: { enable_thinking: true } thoughts block shown, correct
Lemonade gpt-oss-20b-GGUF low chat_template_kwargs: { reasoning_effort: "low" } short thinking, correct

In every case the header, metrics line, and DB row all show the same effort that was sent.

Provider-side confirmation that the levels are honored (direct SDK calls, same prompt):

  • OpenAI gpt-5.1: reasoning tokens 0 (off) → 6 (low) → 197 (high)
  • Anthropic claude-sonnet-5: output tokens 42 (low) → 280 (max)
  • Gemini 3.5-flash: thinking tokens 0 (minimal) → 1429 (high); 2.5-pro 1327 (low) → 2792 (high)
  • DeepSeek: reasoning_content absent (off) / present (on)
  • Ollama qwen3:8b: no thinking (off) / 4693 chars (on)
  • LM Studio: 0 reasoning tokens (off) / 1198 (default)
  • Lemonade gpt-oss-20b: 100 thinking chars (low) → 1694 (high); Qwen3-4B 0 (off) / 5612 (on)

Also verified:

  • system-wide default set on the LLM preference page for all 7 providers, picked up by workspaces with no override
  • workspace override beats the global default; clearing it falls back
  • a stored effort the current model does not support (eg: low on claude-sonnet-4-5, high on gpt-4.1-mini, high on llama3.2) is dropped and the chat still succeeds
  • agent chats for all 7 providers

@shatfield4
shatfield4 marked this pull request as ready for review September 16, 2026 03:45
@timothycarambat

timothycarambat commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

fixes applied below:

1. Model-router workspaces couldn't undo a stored effort
GET /workspace/:slug/llm-capabilities built the connector with getLLMProvider, which throws for anythingllm-router. The picker was hidden, but the chat path still applied the stored effort to whichever routed model listed it, so a thread with "high" set before a switch to the router kept sending it with no way to clear it. There is now a usesModelRouter(workspace) check in the capabilities endpoint, streamChatWithWorkspace (which also covers the effort handed to agents), and AgentHandler#reasoningEffortForRoute (covers mid-session updates). Routed workspaces show no picker and never apply an effort.

2. An offline upgrade wiped the pricing cache
#loadFromDisk read the pricing and reasoning files in one try. After an upgrade the reasoning file doesn't exist, so pricing was nulled too, and air-gapped installs lost cost data. The reasoning file is now read on its own. Pricing stays loaded, and the etag is dropped so the next refresh is a full GET. A new test fails on the previous code.

3. Ollama levels come from Ollama's own API
/api/show returns thinking.values per model, so that list is used now. The gpt-oss name check is kept as a fallback for Ollama versions that don't return it.

4. Sources
Each provider's case in reasoningParams now links to its docs. Two are still unconfirmed:

  • LM Studio's per-model allowed_options is documented, but sending "on" as reasoning_effort: "medium" on the OpenAI-compatible endpoint isn't.
  • Lemonade's chat_template_kwargs depends on the llama.cpp backend.

Both fail safely, since "Provider default" clears the choice.

5. Smaller fixes

  • The home page now keeps its choice under its own draft session, so it no longer writes to the workspace's default chat session.
  • Moved getProviderForConfig's JSDoc back above it.
  • Fixed the "system default" wording in #updateReasoningEffort.

shatfield4 and others added 2 commits October 1, 2026 10:24
- Add a reasoning effort picker to the chat prompt input, scoped per chat
  session (thread or workspace default chat) and kept in the browser only.
- Validate the session's effort against the model's live capabilities
  before every request so unsupported levels are never sent.
- Read cloud reasoning levels from models.dev and local provider levels
  from each provider's model API.
- Map efforts to each provider's request format for chat and agent
  providers, and apply effort changes mid agent session.
- Show OpenAI reasoning summaries and Gemini thoughts as thinking.
…line

- Never apply a session reasoning effort on model-router workspaces. The
  picker is hidden there, so a stored effort could not be seen or cleared.
- Keep the pricing cache when the reasoning cache is missing, so an
  offline upgrade does not lose cost data.
- Read Ollama's per-model think values from /api/show, falling back to
  the model family on older Ollama versions.
- Keep the home page's picker choice under its own draft session instead
  of the workspace's default chat.
- Link each provider's reasoning wire format to its documentation and fix
  stale JSDoc.
@timothycarambat
timothycarambat force-pushed the feat/reasoning-controls branch from ba3a60a to b905294 Compare October 1, 2026 17:26
- Read OpenRouter reasoning levels from models.dev: effort values, with
  "none" as off, and on/off for toggle models. Models with only a token
  budget, or no options (eg: mandatory thinking), get no controls.
- Send the effort with OpenRouter's unified `reasoning` param on chat and
  agent requests, keeping the legacy `include_reasoning` flag when no
  effort is set.
@timothycarambat
timothycarambat merged commit d4e1c6b into master Oct 1, 2026
7 checks passed
@timothycarambat
timothycarambat deleted the feat/reasoning-controls branch October 1, 2026 18:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants