Skip to content

test: CUDA memory pool fallback in python backend - #8992

Merged
mattwittwer merged 4 commits into
mainfrom
mwittwer/python_backend_pinned_fallback_test
Oct 7, 2026
Merged

mattwittwer merged 4 commits into
mainfrom
mwittwer/python_backend_pinned_fallback_test

Conversation

@mattwittwer

@mattwittwer mattwittwer commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

What does the PR do?

Adds regression coverage for silent output corruption in the Python backend when the CUDA memory pool is exhausted and GPU output tensors fall back to pinned host memory (#7148). The runtime fix is python_backend#457.

The existing IOTest.test_ensemble_io already drives the ensemble_io pipeline of three chained dlpack_io_identity Python models, with per-request flags choosing which stage emits its output as a GPU tensor, and asserts exact equality against a 1000 x FP32 input. This change re-runs that test with --cuda-memory-pool-byte-size=0:1024, so every 4000-byte GPU output overflows the pool and the ensemble's response allocator falls back to pinned memory for each of them. The block also fails if the server log does not contain the core's falling back to pinned system memory warning, so it cannot pass without exercising the fallback path. No new models or Python code.

Checklist

  • PR title reflects the change and is of format <commit_type>: <Title>
  • Changes are described in the pull request.
  • Related issues are referenced.
  • Populated github labels field
  • Added test plan and verified test passes.
  • Verified that the PR passes existing CI.
  • Verified copyright is correct on all changed files.
  • Added succinct git squash message before merging.
  • All template sections are filled out.
  • Optional: Additional screenshots for behavior/output changes with before/after.

Commit Type:

  • test

Related PRs:

Where should the reviewer start?

  • qa/L0_backend_python/io/test.sh — the new block IOTest.test_ensemble_io with GPU outputs falling back to pinned memory: model setup mirrors the existing default trial, SERVER_ARGS adds the 1024-byte CUDA pool, and the post-run grep on the server log guards against a vacuous pass.

Test plan:

  • CI Pipeline ID: [71702518]

Caveats:

Background

Related Issues:

@mattwittwer mattwittwer self-assigned this Oct 1, 2026
@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Low risk] Refactors test setup script for Python backend tests.

The PR appears safe to merge based on the reviewed changes.

Findings

  1. P1 Cleanup failure skips remaining checks ▶

Summary

The PR adds CUDA memory-pool fallback coverage for both standard and decoupled Python ensemble IO, using separate server logs and checking that the fallback warning appears.

  • Refactors ensemble model setup so both trials can be rerun with a 1024-byte CUDA pool.
  • Skips test and cleanup commands when a fallback-trial server fails to start.

Reviews (3) · Last reviewed commit: "Merge branch 'mwittwer/python_backend_pi..."

Comment thread qa/L0_backend_python/io/test.sh Outdated
Comment thread qa/L0_backend_python/io/test.sh Outdated
@mattwittwer mattwittwer changed the title draft: test: CUDA memory pool fallback in python backend test: CUDA memory pool fallback in python backend Oct 5, 2026
@mattwittwer
mattwittwer merged commit 7c6031b into main Oct 7, 2026
4 checks passed
@mattwittwer
mattwittwer deleted the mwittwer/python_backend_pinned_fallback_test branch October 7, 2026 19:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants