Skip to content

fix(llmcore): responses payload parity - temperature passthrough, ultra effort (closes #813 A-4/A-5) - #822

Open
louisss1016 wants to merge 1 commit into
lsdefine:mainfrom
louisss1016:fix/issue-813-responses-payload-parity
Open

louisss1016 wants to merge 1 commit into
lsdefine:mainfrom
louisss1016:fix/issue-813-responses-payload-parity

Conversation

@louisss1016

Copy link
Copy Markdown

Fixes the two origin-native llmcore.py payload bugs listed as A-4/A-5 in #813 (root-caused by @GiftedScout — analysis credited, co-author on the commit).

Bug 1 — responses mode drops temperature

_openai_stream's chat_completions branch (and ClaudeSession.raw_ask) both gate on if temperature != 1, but the responses branch never wrote the field: any temperature configured in mykey.py silently no-op'd for every Responses-API endpoint (DeepSeek etc. via api_mode: responses). The official Responses API supports temperature (default 1.0).

Fix: same if temperature != 1 guard appended after the reasoning / max_output_tokens lines. The kimi/moonshot force-to-1 (line 519) and MiniMax clamp (line 520) happen before branch selection, so those endpoints are unaffected.

Bug 2 — reasoning_effort: ultra rejected by the whitelist

BaseSession._enum's valid set ends at 'max'; GPT-5.6's ultra was dropped with a single [WARN] Invalid reasoning_effort 'ultra', ignored. and the payload silently fell back to the endpoint default. Fix: one-line whitelist addition in BaseSession.__init__ (shared by every session type, so this covers chat + responses + native alike).

Note: the Claude-only output_config.effort mapping ({'low','medium','high','xhigh'->'max','max'->'max'}) still doesn't map ultra — it warns there instead. That path is Claude-specific and out of scope; flag in case @cuipengcx90 wants it addressed.

Tests — frontends/tests/test_responses_payload_parity.py (7 tests): monkeypatch llmcore.requests.post and drive a real NativeOAISession.raw_ask with a canned SSE stream — no network, no credentials:

  • temperature present (0.5) in responses payload when configured
  • temperature omitted at default 1 (payload stays identical for most configs)
  • temperature still omitted for kimi model override (forced to 1)
  • ultra survives _enum → sess.reasoning_effort == 'ultra' and payload['reasoning']['effort'] == 'ultra'
  • high unchanged; bogus still dropped (WARN path intact)
  • chat_completions branch shape pinned: flat reasoning_effort key, max_tokens key for non-gpt-5 models

Full suite: 279 passed; the 3 failures (release_qualification, symlink ×2) are the same pre-existing Windows-env baseline.

The fake response implements the context-manager protocol because upstream now calls requests.post under with (TTFT/abort support) — added after @GiftedScout's fork point, so their standalone test script needed adapting.

…ra effort

Two silent-config bugs in the OpenAI Responses branch (issue lsdefine#813, A-4/A-5,
root-caused by GiftedScout):

1. temperature: the chat_completions branch and ClaudeSession.raw_ask both
   send the user-configured value behind `if temperature != 1`; the responses
   branch never did, so temperature in mykey.py silently no-op'd for every
   responses endpoint. Official Responses API supports the field (default
   1.0), so the chat-branch guard is reused verbatim - the kimi/moonshot
   force-to-1 and MiniMax clamp upstream of the branch still apply.
2. reasoning_effort 'ultra': BaseSession._enum's whitelist ended at 'max',
   so GPT-5.6's 'ultra' was dropped with a single WARN line and the payload
   fell back to the endpoint default. One-line whitelist addition; the
   Claude-only output_config.effort mapping still warns on ultra there.

Tests: frontends/tests/test_responses_payload_parity.py - monkeypatches
requests.post to capture the real payload through a real
NativeOAISession.raw_ask (canned SSE, no network/credentials):
temperature present when configured / omitted at default 1 / omitted under
model overrides, ultra survives _enum and reaches payload, high unchanged,
bogus still dropped, chat_completions branch shape pinned as regression
guard. Full suite: 279 passed, same 3 pre-existing Windows-env failures.

Co-Authored-By: GiftedScout <pushuai66@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant