Skip to content

fix(core): forward task images into the first LLM turn (multimodal passthrough) - #828

Open
louisss1016 wants to merge 1 commit into
lsdefine:mainfrom
louisss1016:fix/issue-813-multimodal-image-passthrough
Open

louisss1016 wants to merge 1 commit into
lsdefine:mainfrom
louisss1016:fix/issue-813-multimodal-image-passthrough

Conversation

@louisss1016

Copy link
Copy Markdown

Part of #813 (item A-1 — multimodal images silently dropped). Follow-up to #822 (A-4/A-5).

Problem

put_task() has always accepted images and every frontend passes them (fsapp downloads Feishu images to disk and forwards the paths; desktop_bridge and tui_v3 forward paths too), but agentmain.run() never handed them to agent_runner_loop — the first user turn stayed plain text and attachments were silently dropped. The desktop bridge papers over this with a backend.ask monkey-patch; fsapp and tui have no such workaround.

Even when image blocks do reach the LLM layer, two more drop points exist:

  1. _to_responses_input() only understood OpenAI-style image_url blocks — Claude-style {"type": "image", "source": {...}} blocks vanished on the Responses API path.
  2. NativeToolClient.chat()'s whitespace filter ([c for c in combined_content if c.get("text", "").strip()]) dropped every block without non-blank text, image blocks included.

Fix

  • agentmain.py: new _multimodal_initial_content() builds the first turn's content blocks (text + base64 image blocks); run() passes them via agent_runner_loop(initial_user_content=...). Gated on NativeToolClient — only those backends understand image blocks (native Claude API takes them as-is; the OAI paths convert them), so other clients keep plain-text behaviour. Non-vision formats (e.g. SVG) degrade to a text path reference, and unreadable files surface as an error text block instead of vanishing silently.
  • llmcore.py: _to_responses_input() converts Claude-style image blocks (base64 and url sources) to input_image, mirroring _msgs_claude2oai().
  • llmcore.py: extract _drop_blank_text_blocks(); only blank text blocks are dropped now, image/tool_result blocks survive.
  • desktop_bridge.py: remove the _patch_chat_for_images monkey-patch — with the core plumbing fixed it would double-inject every image. Attachments flow through put_task() instead. The guard tests in test_bridge_submit.py are repointed at the core mechanism (test_core_forwards_images_to_agent_loop, test_run_agent_turn_does_not_monkeypatch_images).

Tests

New frontends/tests/test_multimodal_image_passthrough.py (21 cases, red before the fix):

  • _to_responses_input: base64/url image blocks -> input_image, OpenAI-style image_url regression guard, assistant role and malformed blocks ignored
  • _drop_blank_text_blocks: blank text dropped, image/tool_result blocks kept
  • NativeToolClient.chat end-to-end: image blocks reach backend.ask, blank text still dropped
  • _multimodal_initial_content (extracted via ast, no import): no images / non-native client -> None, text+image blocks with verified base64 bytes, unknown extension -> image/png, SVG -> text reference, unreadable path -> error text, dict-shaped entries
  • run() wiring: agent_runner_loop receives initial_user_content fed from task.get("images")

Full suite: 294 passed, with the same 3 pre-existing environmental failures as the untouched baseline (272 passed + symlink x2 + release packaging) — zero regressions, +22 new passing tests.

…A-1)

put_task() has always accepted images and every frontend passes them
(fsapp downloads Feishu images to disk, desktop_bridge/tui_v3 forward
paths), but run() never handed them to agent_runner_loop — the first
user turn stayed plain text and attachments were silently dropped.

- agentmain: new _multimodal_initial_content() builds Claude-style
  content blocks (text + base64 image blocks); run() passes them via
  agent_runner_loop(initial_user_content=...). Gated on NativeToolClient
  so other clients keep plain-text behaviour. Non-vision formats (e.g.
  SVG) degrade to a text path reference, and unreadable files surface
  as an error text block instead of vanishing.
- llmcore: _to_responses_input() now converts Claude-style image blocks
  to input_image. Previously only OpenAI-style image_url blocks
  survived, so attachments died on the Responses API path even when
  they reached the converter.
- llmcore: extract _drop_blank_text_blocks(). The old inline filter
  dropped every block without non-blank text — image blocks included —
  so NativeToolClient.chat() stripped attachments before the backend
  ever saw them.
- desktop_bridge: remove the _patch_chat_for_images monkey-patch. It
  existed to work around the missing core plumbing and would now
  double-inject every image; attachments flow through put_task instead.
  Its guard tests are repointed at the core mechanism.

Tests: frontends/tests/test_multimodal_image_passthrough.py (21 cases:
responses converter image/url/malformed/assistant branches, blank-text
filter keeping image/tool_result blocks, NativeToolClient.chat end-to-end
passthrough, helper extraction via ast, run() wiring). Full suite:
294 passed with the same 3 pre-existing environmental failures as the
untouched baseline (272 passed + symlink x2, release packaging).

Refs lsdefine#813 (item A-1)

Co-Authored-By: GiftedScout <GiftedScout@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant