Repository navigation
test: CUDA memory pool fallback in python backend - #8992
Merged
Merged
Conversation
…s back to pinned memory
|
….com:triton-inference-server/server into mwittwer/python_backend_pinned_fallback_test
Merged
7 of 11 tasks
whoisj
approved these changes
Oct 7, 2026
Vinya567
approved these changes
Oct 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does the PR do?
Adds regression coverage for silent output corruption in the Python backend when the CUDA memory pool is exhausted and GPU output tensors fall back to pinned host memory (#7148). The runtime fix is
python_backend#457.The existing
IOTest.test_ensemble_ioalready drives theensemble_iopipeline of three chaineddlpack_io_identityPython models, with per-request flags choosing which stage emits its output as a GPU tensor, and asserts exact equality against a 1000 x FP32 input. This change re-runs that test with--cuda-memory-pool-byte-size=0:1024, so every 4000-byte GPU output overflows the pool and the ensemble's response allocator falls back to pinned memory for each of them. The block also fails if the server log does not contain the core'sfalling back to pinned system memorywarning, so it cannot pass without exercising the fallback path. No new models or Python code.Checklist
<commit_type>: <Title>Commit Type:
Related PRs:
test guards; the new block fails until it is in the test container)
Where should the reviewer start?
qa/L0_backend_python/io/test.sh— the new blockIOTest.test_ensemble_io with GPU outputs falling back to pinned memory: model setup mirrors the existingdefaulttrial,SERVER_ARGSadds the 1024-byte CUDA pool, and the post-rungrepon the server log guards against a vacuous pass.Test plan:
Caveats:
Background
Related Issues: