Summary
A float column that contains both +inf and -inf writes a different file on x86_64 than on
aarch64, for the same input and the same build profile. The difference is exactly one bit: the sign
of a NaN stored as the column's Sum statistic.
Still present on develop (checked at the current tip; sum_v2 has the same property).
Root cause
sum_float_all (vortex-array/src/aggregate_fn/fns/sum/primitive.rs) steps over NaN inputs when
skip_nans is set, so the two infinities meet in the accumulator and inf + -inf is evaluated. That
is an IEEE 754 invalid operation, and IEEE does not specify the resulting bit pattern. The targets
disagree:
| target |
inf + -inf |
| x86_64 (SSE) |
0xfff8_0000_0000_0000 — sign bit set |
| aarch64 |
0x7ff8_0000_0000_0000 — sign bit clear |
Stat::Sum is in PRUNING_STATS (documented as the stats "we want to ensure [are] computed when
compressing/writing") and StatsSet::write_flatbuffer serialises it, so that platform-chosen value
lands in the file.
Note the NaN-skipping is what exposes this. Without it the accumulator would be NaN from the first
NaN element and would propagate that element's payload, which both architectures do identically.
Skipping is what lets the infinities meet and generate a fresh default NaN.
Reproducer
One f64 column, 7 rows:
0x7ff8_0000_dead_beef (NaN with a payload)
-0.0
0.0
f64::INFINITY
f64::NEG_INFINITY
f64::MIN_POSITIVE
0x000f_ffff_ffff_ffff (largest subnormal)
Written on both architectures in one CI matrix, same pinned rustc, same debug profile, so the target
architecture is the only variable. Bisecting cumulative prefixes of that column:
| case |
length (both) |
result |
| ordinary finite control |
2564 |
byte-identical |
| each of the 7 values alone |
2204 / 2548 |
byte-identical |
prefixes 1–4 (+inf present, -inf absent) |
2204–2572 |
byte-identical |
prefix 5 (adds -inf) |
2580 |
1 byte, 1 bit differs |
| prefix 6, 7, full column |
2588 |
1 byte, 1 bit differs |
The differing byte is the top byte of an 8-byte little-endian field:
x86_64 : … 00 00 00 00 00 00 f8 ff … -> fff8000000000000
aarch64: … 00 00 00 00 00 00 f8 7f … -> 7ff8000000000000
So the whole divergence is one bit. The trigger is precisely "both infinities in one column",
which is what makes inf + -inf reachable: prefix 4 already has +inf, and prefix 5 is the first to
add -inf.
Measured hardware semantics from the same build, for the record:
x86_64 inf-inf = inf/inf = 0/0 = 0*inf = fff8000000000000
aarch64 inf-inf = inf/inf = 0/0 = 0*inf = 7ff8000000000000
both NaN payload propagation is IDENTICAL (0x7ff80000deadbeef survives +0.0 and *1.0),
and f64::min/max correctly ignore NaN on both.
Only the default NaN generated by an invalid operation differs.
Why it matters
We content-address written files, so the hash of the bytes is the artifact's identity. This makes the
identity of the same logical table depend on which machine wrote it, and anything verifying a digest
across a heterogeneous fleet sees corruption where there is none. It is also profile-independent, so
it is not caught by building everything one way.
Fix
PR to follow: report any NaN sum as the canonical quiet NaN, in both sum and sum_v2 (each
finalises floats through its own path). A NaN sum carries no payload information, so nothing is lost,
and it makes the computed NaN agree with the f64::NAN that sum_v2 already writes explicitly on
its non-skip_nans poisoning path.
Alternatives, if you would prefer one: record Sum as absent/inexact when the column contains both
infinities, or track ±inf counts separately from the finite sum. I went with canonicalisation as the
smallest change that makes the stat a function of the data — but is_saturated already treats a NaN
float sum as terminal, so the codebase currently treats it as a real value, and making it absent
would be the larger semantic change.
Downstream tracking: vig-os/tessera#472.
Summary
A float column that contains both
+infand-infwrites a different file on x86_64 than onaarch64, for the same input and the same build profile. The difference is exactly one bit: the sign
of a NaN stored as the column's
Sumstatistic.Still present on
develop(checked at the current tip;sum_v2has the same property).Root cause
sum_float_all(vortex-array/src/aggregate_fn/fns/sum/primitive.rs) steps over NaN inputs whenskip_nansis set, so the two infinities meet in the accumulator andinf + -infis evaluated. Thatis an IEEE 754 invalid operation, and IEEE does not specify the resulting bit pattern. The targets
disagree:
inf + -inf0xfff8_0000_0000_0000— sign bit set0x7ff8_0000_0000_0000— sign bit clearStat::Sumis inPRUNING_STATS(documented as the stats "we want to ensure [are] computed whencompressing/writing") and
StatsSet::write_flatbufferserialises it, so that platform-chosen valuelands in the file.
Note the NaN-skipping is what exposes this. Without it the accumulator would be NaN from the first
NaN element and would propagate that element's payload, which both architectures do identically.
Skipping is what lets the infinities meet and generate a fresh default NaN.
Reproducer
One f64 column, 7 rows:
Written on both architectures in one CI matrix, same pinned rustc, same debug profile, so the target
architecture is the only variable. Bisecting cumulative prefixes of that column:
+infpresent,-infabsent)-inf)The differing byte is the top byte of an 8-byte little-endian field:
So the whole divergence is one bit. The trigger is precisely "both infinities in one column",
which is what makes
inf + -infreachable: prefix 4 already has+inf, and prefix 5 is the first toadd
-inf.Measured hardware semantics from the same build, for the record:
Only the default NaN generated by an invalid operation differs.
Why it matters
We content-address written files, so the hash of the bytes is the artifact's identity. This makes the
identity of the same logical table depend on which machine wrote it, and anything verifying a digest
across a heterogeneous fleet sees corruption where there is none. It is also profile-independent, so
it is not caught by building everything one way.
Fix
PR to follow: report any NaN sum as the canonical quiet NaN, in both
sumandsum_v2(eachfinalises floats through its own path). A NaN sum carries no payload information, so nothing is lost,
and it makes the computed NaN agree with the
f64::NANthatsum_v2already writes explicitly onits non-
skip_nanspoisoning path.Alternatives, if you would prefer one: record
Sumas absent/inexact when the column contains bothinfinities, or track
±infcounts separately from the finite sum. I went with canonicalisation as thesmallest change that makes the stat a function of the data — but
is_saturatedalready treats a NaNfloat sum as terminal, so the codebase currently treats it as a real value, and making it absent
would be the larger semantic change.
Downstream tracking: vig-os/tessera#472.