Skip to content

fix: Preserve generate model error status - #9002

Open
amarrtech wants to merge 2 commits into
triton-inference-server:mainfrom
amarrtech:fix/http-generate-error-status
Open

amarrtech wants to merge 2 commits into
triton-inference-server:mainfrom
amarrtech:fix/http-generate-error-status

Conversation

@amarrtech

Copy link
Copy Markdown

What does the PR do?

Preserves the original Triton model error status for the HTTP /generate and /generate_stream endpoints.

GenerateRequestClass::InferResponseComplete previously passed err to AddErrorJson, which serializes and deletes it, and then called HttpCodeFromError(err). Capturing the HTTP code before serialization removes that use-after-free and restores the intended mappings, including INTERNAL to 500 and UNAVAILABLE to 503.

The L0 HTTP regression drives both endpoints through first-response INTERNAL and UNAVAILABLE failures using the existing mock_llm model.

Checklist

  • I have read the Contribution guidelines and signed the Contributor License Agreement — guidelines read; NVIDIA CLA confirmation is still pending.
  • PR title reflects the change and is of format <commit_type>: <Title>
  • Changes are described in the pull request.
  • Related issues are referenced.
  • Populated github labels field — external contributor permissions may require a maintainer to apply the label.
  • Added test plan and verified test passes — the regression is wired into L0 HTTP, but the packaged Linux Triton test environment was unavailable locally.
  • Verified that the PR passes existing CI — awaiting hosted checks.
  • I ran pre-commit locally (pre-commit install, pre-commit run --all) — all hooks passed on the four changed files; the repository-wide suite was not run.
  • Verified copyright is correct on all changed files.
  • Added succinct git squash message before merging ref.
  • All template sections are filled out.
  • Optional: Additional screenshots for behavior/output changes with before/after.

Commit Type:

  • build
  • ci
  • docs
  • feat
  • fix
  • perf
  • refactor
  • revert
  • style
  • test

Related PRs:

None.

Where should the reviewer start?

src/http_server.cc, where HttpCodeFromError(err) now runs before AddErrorJson(err) deletes the error. The regression entry point is qa/L0_http/generate_error_status_test.py.

Test plan:

  • uvx --from pre-commit pre-commit run --show-diff-on-failure --color=never --files src/http_server.cc qa/L0_http/test.sh qa/L0_http/generate_error_status_test.py qa/python_models/generate_models/mock_llm/1/model.py — passed all hooks.
  • python3 -m py_compile qa/L0_http/generate_error_status_test.py qa/python_models/generate_models/mock_llm/1/model.py — passed.
  • git diff --check — passed.
  • qa/L0_http/test.sh — not run locally; it requires a packaged Linux Triton server and Python backend. The new regression is invoked by this suite and covers 500/503 mappings on both generate endpoints.

Caveats:

The full L0 HTTP runtime test requires Triton's packaged Linux QA environment, which was unavailable from this macOS source checkout.

Background

Gateways and clients use HTTP status codes to distinguish caller errors from model-engine failures. Returning 400 for an unhealthy engine can suppress appropriate retry, failover, and health behavior and obscures the operator diagnosis.

OpenAI Codex assisted with the implementation and regression test; the commit records that assistance.

Related Issues: (use one of the action keywords Closes / Fixes / Resolves / Relates to)

Capture the HTTP response code before serializing and deleting the Triton error. Add generate and generate_stream coverage for INTERNAL and UNAVAILABLE model failures.

Assisted-by: OpenAI Codex
Signed-off-by: Amrinder Randhawa <272048731+amarrtech@users.noreply.github.com>
@greptile-apps

greptile-apps Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium risk] Fixes error status handling in the generate endpoint.

The PR appears safe to merge; the previous timeout issue is fixed and no new blocking issue was found.

What we checked:

  • Timeout rejects a valid test response: These requests set REPETITION=0 and omit DELAY. The model skips its delayed response loop and sends the final error immediately.

Summary

Preserves model error status on /generate and /generate_stream by reading the HTTP code before AddErrorJson deletes the error.

  • Adds regression checks for 500 and 503 responses on both endpoints.
  • Adds those checks to the L0 HTTP suite.
  • Bounds each regression request with a 30-second timeout, addressing the previous finding.
  • No new actionable issues found. The packaged L0 HTTP suite was not run during this review.

Reviews (2) · Last reviewed commit: "test: bound generate error requests" · Reviewed by Greptile

Comment thread qa/L0_http/generate_error_status_test.py
Assisted-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Amrinder Randhawa <272048731+amarrtech@users.noreply.github.com>
@greptile-apps

greptile-apps Bot commented Oct 8, 2026

Copy link
Copy Markdown

Want your agent to iterate on Greptile's feedback? Try greploops.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

HTTP generate/generate_stream return 400 for all model errors (error freed before HttpCodeFromError)

1 participant