Repository navigation
fix(llmcore): responses payload parity - temperature passthrough, ultra effort (closes #813 A-4/A-5) - #822
Open
louisss1016 wants to merge 1 commit into
Conversation
…ra effort Two silent-config bugs in the OpenAI Responses branch (issue lsdefine#813, A-4/A-5, root-caused by GiftedScout): 1. temperature: the chat_completions branch and ClaudeSession.raw_ask both send the user-configured value behind `if temperature != 1`; the responses branch never did, so temperature in mykey.py silently no-op'd for every responses endpoint. Official Responses API supports the field (default 1.0), so the chat-branch guard is reused verbatim - the kimi/moonshot force-to-1 and MiniMax clamp upstream of the branch still apply. 2. reasoning_effort 'ultra': BaseSession._enum's whitelist ended at 'max', so GPT-5.6's 'ultra' was dropped with a single WARN line and the payload fell back to the endpoint default. One-line whitelist addition; the Claude-only output_config.effort mapping still warns on ultra there. Tests: frontends/tests/test_responses_payload_parity.py - monkeypatches requests.post to capture the real payload through a real NativeOAISession.raw_ask (canned SSE, no network/credentials): temperature present when configured / omitted at default 1 / omitted under model overrides, ultra survives _enum and reaches payload, high unchanged, bogus still dropped, chat_completions branch shape pinned as regression guard. Full suite: 279 passed, same 3 pre-existing Windows-env failures. Co-Authored-By: GiftedScout <pushuai66@gmail.com>
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the two origin-native
llmcore.pypayload bugs listed as A-4/A-5 in #813 (root-caused by @GiftedScout — analysis credited, co-author on the commit).Bug 1 — responses mode drops
temperature_openai_stream's chat_completions branch (andClaudeSession.raw_ask) both gate onif temperature != 1, but the responses branch never wrote the field: any temperature configured in mykey.py silently no-op'd for every Responses-API endpoint (DeepSeek etc. viaapi_mode: responses). The official Responses API supportstemperature(default 1.0).Fix: same
if temperature != 1guard appended after the reasoning / max_output_tokens lines. The kimi/moonshot force-to-1 (line 519) and MiniMax clamp (line 520) happen before branch selection, so those endpoints are unaffected.Bug 2 —
reasoning_effort: ultrarejected by the whitelistBaseSession._enum's valid set ends at'max'; GPT-5.6'sultrawas dropped with a single[WARN] Invalid reasoning_effort 'ultra', ignored.and the payload silently fell back to the endpoint default. Fix: one-line whitelist addition inBaseSession.__init__(shared by every session type, so this covers chat + responses + native alike).Note: the Claude-only
output_config.effortmapping ({'low','medium','high','xhigh'->'max','max'->'max'}) still doesn't mapultra— it warns there instead. That path is Claude-specific and out of scope; flag in case @cuipengcx90 wants it addressed.Tests —
frontends/tests/test_responses_payload_parity.py(7 tests): monkeypatchllmcore.requests.postand drive a realNativeOAISession.raw_askwith a canned SSE stream — no network, no credentials:ultrasurvives_enum→sess.reasoning_effort == 'ultra'andpayload['reasoning']['effort'] == 'ultra'highunchanged;bogusstill dropped (WARN path intact)reasoning_effortkey,max_tokenskey for non-gpt-5 modelsFull suite: 279 passed; the 3 failures (release_qualification, symlink ×2) are the same pre-existing Windows-env baseline.
The fake response implements the context-manager protocol because upstream now calls
requests.postunderwith(TTFT/abort support) — added after @GiftedScout's fork point, so their standalone test script needed adapting.