Conversation
remer
force-pushed
the
fix/cuda-glibc-compat
branch
from
September 2, 2026 20:23
0ee1b0d to
6963514
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Make Krasis's Linux CUDA builds tolerate the CUDA 13.0/13.1 declaration
collision with newer glibc without patching the installed CUDA toolkit or
changing the system compiler.
-U_GNU_SOURCE -D_DEFAULT_SOURCEthrough the sharedbuild.rsNVCCargument path;
FlashAttention package sidecars;
optional libstdc++ pthread clock-wait paths made unavailable when
_GNU_SOURCEis hidden; andcompatibility-source change.
Windows behavior is unchanged. An explicitly selected
KRASIS_NVCC_CCBINisstill preserved.
Observed failures
With CUDA 13.1 and glibc 2.43, normal CUDA compilation failed with errors
equivalent to:
After the original compatibility flags were exercised through the release
sidecar builder, FlashAttention exposed the corresponding libstdc++ half of the
contract: its optional timed-mutex paths referenced GNU pthread clock APIs that
were intentionally no longer declared. The compatibility header handles both
sides in the affected translation unit.
RC10 validation (2026-09-02)
v1.0.21-rc.10(
716bf316fbc001c1f61ad607362e05c1cb323c50).git diff --checkand Rust formatting checks passed.for head dimensions 64, 96, 128, 192, and 256, causal and non-causal.
manifest input hashes, artifact hashes, and build IDs all passed the
builder's post-build contract.
and revalidated, proving the new header participates in cache identity.
3.13; both packaged sidecars loaded from that wheel and 29 wheel-backed
DeepSeek Vision/model-configuration tests passed.
The DeepSeek-only canary does not require Krasis's separate Flash Linear
Attention sidecars, so this validation does not claim a full multi-model
release-wheel matrix.
Compatibility scope
The observed CUDA 13.1/new-glibc combination is outside CUDA 13.1's validated
distribution matrix, so this is a portability improvement rather than a claim
that Krasis is defective on a supported CUDA configuration. NVIDIA's
cuda-samplesPR #403 documents the same feature-macro workaround:NVIDIA/cuda-samples#403
No credentials, account names, hostnames, network addresses, private paths, or
raw request/generated content are included in this change or PR body.
RC11 alignment — 2026-09-08
Merged upstream
v1.0.21-rc.11(9d25fb3a490499f1be78952079d361acfdb5c05c) without rewriting thisbranch's history. The feature diff is unchanged; RC11 contributes its 13
context-default/launcher/manager/version/test files. Previous head preserved
at
remer/krasis:preserve/pre-rc11-20260908/pr31.The combined
integration/rc11-ai-boxtreed8656f9577a7d469a7b0e2e987b61b14e258e1d5builds a fresh release wheel.Its targeted Rust suites pass: D-Spark 30/30, chat template 16/16, server 31/31.
Python model/configuration/vision/cache-gate contracts pass. One launcher parser
case assumes two visible GPUs; it passes with an isolated two-GPU inventory
fixture (parser-only, not a claim of multi-GPU runtime testing).
Live RC11 Vision qualification is pending. Historical RC6/RC10 throughput and
quality measurements are not RC11 Vision results. Session-cache operational
and D-Spark promotion gates remain live-test obligations.