Skip to content

fix(zarr-metadata): close the four gaps a reader of foreign stores hits - #381

Open
d-v-b wants to merge 4 commits into
feat/zarr-metadata-validate-at-constructionfrom
fix/zarr-metadata-foreign-document-gaps
Open

d-v-b wants to merge 4 commits into
feat/zarr-metadata-validate-at-constructionfrom
fix/zarr-metadata-foreign-document-gaps

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 29, 2026

Copy link
Copy Markdown
Owner

🤖 AI text below 🤖

A member a closed object does not declare is an unknown_key wherever
it sits, as ProblemKind defines one: in a v2 .zgroup, a .zmetadata
envelope, a group's consolidated_metadata and a metadata field's
envelope, where it was an invalid_value, as it already was inside a v3
configuration. So a reader that tolerates a key another writer added,
such as netCDF-C's _nczarr_* keys in a .zgroup, can filter by kind.
A definition's check and judge give a configuration back when a
field it holds has a stray member, which they report and leave out, as
the checker does a key a closed TypedDict does not declare.

A zarr.json that says no node type the spec defines is judged by its
zarr_format too, so a document of another format says it is not v3.
The root zarr.json of zarr-python 2's draft of v3, whose zarr_format
is a URL, was reported only as missing its node_type; it now also has
an invalid_type problem at zarr_format, and a v2 document read as a
zarr.json an invalid_value there.

A group's consolidated metadata is judged as the hierarchy below the
group, the group its root: zarr-python keeps the document of the node at
/a/b at the key a/b, and the documents and the group make a tree in
which only groups hold nodes and each node's parent is held. A key that
is no node's path below the group, such as "", "__a" or "a/../b",
and a document below an array's were read without a problem, and so was
a document whose group is missing, which zarr-python fails to read. Each
is a problem now, a missing group a missing_key, and
ZarrV3ConsolidatedMetadata refuses them when it is built.

Two v3 array models are equal when their fill values are one value of the
data type, however each is spelled. "NaN" and "0x7fc00000" were two
float32 fill values, and 0.0 and -0.0, which are two values of a
float, were one, since == compared the JSON. A fill value whose data
type is not interpreted -- one nothing in scope claims, or a v2 model's
-- is compared as before.

NodeName and NodePath, in zarr_metadata.v3, are the names and paths
of the nodes of a v3 hierarchy, modelled on zarrs' types of those names,
and zarr_metadata.model judges a string by the spec's rules for them:
validate_node_name_v3, is_node_name_v3 and parse_node_name_v3, and
their node_path twins.

A data type's definition says which spelling of a fill value is its
value's own: DataTypeDefinition.fill_value_canonical, and
canonical_fill_value(data_type, value) in zarr_metadata.v3.definition
spells one, or gives UNSET for a fill value with a problem. The float
types read a number as a float64, as JSON parsers and numpy do, and
spell a value by its shortest number, a named value, or the bits of a
NaN the spec does not name; the complex types spell each part so; the
numpy time types spell -2**63 as "NaT"; bytes spells its bytes as
base64; and a struct spells each field's fill value by that field's
type.

What changes for a caller

before now
NCZarr's _nczarr_group in a .zgroup invalid_value unknown_key, "unexpected key '_nczarr_group'"
judge/check of a configuration whose nested field has a stray member None the configuration, the member reported and left out
zarr-python 2's draft-v3 root zarr.json node_type missing also zarr_format: invalid_type, "expected 3, got" the URL
consolidated key "__a", "", "a/../b" or "/a" read invalid_value at the entry
a document at a/b where a is an array read invalid_value at a/b: 'expected a node below a group, got "/a/b", below the array "/a"'
a document at a/b/c with no a/b read (zarr-python raises KeyError) missing_key at a/b
float32 fill "NaN" vs "0x7fc00000" two models one
float32 fill 0.0 vs -0.0 one model two

How it works

  • Kinds. unexpected_keys and the envelope check report a stray member as unknown_key. judge puts each nested field back with only the members its envelope declares, as resolve already did.

  • Format. When the node-type dispatch cannot follow a document, _node_type judges its zarr_format too, for the reader and the validator alike.

  • Hierarchy. A private hierarchy_problems(nodes) in v3/_hierarchy.py takes the node type of each node by its node path. It judges them as the spec's tree:

    • each key is a node path;
    • no node sits below an array;
    • each node's parent is present, the root's included.

    It knows nothing of consolidated metadata, so a future model of a whole hierarchy can use it as is. Consolidated metadata is one caller: the group is the root /, and the key a/b is the node at /a/b. A key's own form is judged where the entry is read, before the entry's problems. The tree is judged once every entry is read.

  • Fill values. DataTypeDefinition.fill_value_canonical is a hook like fill_value_rules, defaulting to "as written". ZarrV3ArrayMetadata.__eq__ compares the canonical spellings as JSON text, and __hash__ agrees with it. The spelling is computed only when models are compared, so reads cost nothing more. For floats:

    • a number is read as a float64 first, as JSON parsers and numpy do;
    • float_bits gives the value's bits in the type;
    • the spelling is the nearest of the shortest decimals that round-trip, which matches numpy's shortest repr.

Verification

  • Tests. 1,804 pass on Python 3.14, and 1,802 plus 2 skipped on 3.11. Each of the four commits passes on its own. pyright, ruff and the strict docs build are clean.

  • Mutation checks. Five deliberate breakages of the float code each fail the tests: the named NaN spelled as bits, the sign lost on either overflow path, the neighbouring shortest decimal not tried, and the signed-zero guard dropped. Two first survived, the sign on the integer overflow path and the signed-zero guard, so the agent added a float16 -0.0 row and a negative integer past the float64 range.

  • Canonical spellings against numpy. All float16 values, every float32 power of two and 200,000 random float32 values: 0 mismatches with numpy's shortest repr.

  • Differential. 481,228 records, compared between feat(zarr-metadata)!: check each model when it is built, and write it as it is #379 and this PR. Every difference is explained by one of the fixes:

    • the kind change: 87,705 v2 group records, 7,466 v3 array records, and more;
    • the zarr_format problem: 21 records;
    • the hierarchy rules: 7,689 records, of which 3,470 turn from valid to invalid, all with synthetic consolidated keys such as a/b without a.

    The fill-value change moves no record.

  • Real documents. 8,808 zarr.json files: the scenarios' stores, test corpora, and documents fetched or regenerated from zarr-python issues. 171 hold consolidated metadata, 121 of them with nested paths and 1,005 entries in all. The only change is the expected one: the two draft-v3 roots now report their zarr_format. No consolidated document is newly refused.

  • Scenarios. The four earlier scenarios replay unchanged. A fifth, a catalog reader of 18 stores that other tools wrote (NCZarr, n5-zarr, IDR, EMBL, earthkit, tensorstore, kerchunk, zarr 2.18 to 3.1), was built for this PR and replayed on it:

    • Its NCZarr workaround is no longer needed.
    • Its draft stores now say they are not v3.
    • Its copy check tells 0.0 from -0.0.
    • Its verdict counts are unchanged.

Reviews

  • roborev, one pass per commit. It found:

    • judge gave a nested field back with its stray member;
    • docs were stale after the zarr_format change;
    • a missing reader/validator agreement test;
    • a misleading "missing group" message above an array;
    • an undocumented treatment of unknown-node documents;
    • an overclaiming README sentence;
    • a pointless float64 digit search.

    All are fixed. It also questioned the zarr.json node-name rule; the pinned spec lists it (core L818-L837).

  • Design review. It found:

    • integers rounded exactly while fractions went through float64, so one number had two values and disagreed with numpy;
    • canonical_fill_value returned None both for a problem and for a JSON null fill value;
    • a __hash__ that disagreed with __eq__;
    • v2 0 vs 0.0 newly unequal;
    • non-shortest spellings at some powers of two, and a signed-zero spelling bug;
    • validators missing the _v3 suffix;
    • problem order;
    • stale anchors and wording.

    All are fixed.

Choices worth a look

  1. The hierarchy check is keyed by the spec's node paths (/a/b), and messages name nodes that way, while a consolidated key is a/b. zarr-python's create_hierarchy keys by relative paths, with "" for the root. If a future hierarchy model uses that form instead, only the key conversion changes.
  2. What is not interpreted keeps Python's ==: a v2 fill value, the fill value of a data type nothing in scope claims, attributes, extra_fields, and an Unclaimed field's configuration. == takes true for 1 and -0.0 for 0.0. Comparing that JSON "as written" or "as JSON Schema compares" is a separate decision, so the README says how it compares today.
  3. .zarray holding attributes stays an invalid_value. The v2 spec lets .zarray hold other keys (readers ignore them), while .zgroup forbids them, so unknown_key would be wrong there.
  4. A problem about a key carries the entry's document as its input, as the package's unknown_key problems already do. pydantic instead locates dict-key errors at (..., key, "[key]"), with the key as input.

Left out

  • A public hierarchy model, which will be the second caller of hierarchy_problems.
  • Comparing uninterpreted JSON by anything but == (choice 2).
  • The scenario's other findings, relayed separately:
    • the model drops NCZarr's .zarray members;
    • a tolerated unknown_key leaves no model;
    • .zmetadata entries are checked only as JSON;
    • consolidated_metadata: null goes unrecorded;
    • the v3 readers take parsed JSON only;
    • attributes in .zarray hides the document's other problems.
Stack

Stacked on #379 (feat/zarr-metadata-validate-at-construction), itself on #378 → #377 → #376 → #374 → #373 → #372 → #371 → #370. Merge after #379. Fragments: 381.bugfix.md, 381.bugfix.1.md, 381.bugfix.2.md, 381.bugfix.3.md, 381.feature.md, 381.feature.1.md. The fragments were numbered for this PR, #381. If the number differs, rename them.

🤖 Generated with Claude Code

…nknown key wherever it sits

`ProblemKind` defines `unknown_key` as a key an object's type does not
declare where the type is closed, and a v3 configuration reported one
so. A v2 `.zgroup`, a `.zmetadata` envelope, a group's
`consolidated_metadata` and a metadata field's envelope reported the same
case as `invalid_value`, so a reader that tolerates what another writer
added -- netCDF-C's `_nczarr_*` keys (zarr-python#2296) -- could not
filter it by kind. Each is now `unknown_key`, with the checker's message,
"unexpected key 'x'". A definition's `check` and `judge` give a
configuration back when a field it holds has a stray member, reported
and left out, as the checker leaves out a key a closed TypedDict does not
declare; a `must_understand` of `false` still refuses it.

Assisted-by: ClaudeCode:claude-opus-5-5
…s in

`read_node_metadata_v3` reads a document by the node type it names, and
one that names none was reported only as missing its `node_type`. A
crawler meeting the root `zarr.json` of zarr-python 2's draft of v3
(zarr-python#2982), whose `zarr_format` is a URL, learned "no node type",
not "not v3". A document the dispatch cannot follow is now judged by its
`zarr_format` as well: missing, or other than 3, is a problem beside the
node type's.

Assisted-by: ClaudeCode:claude-opus-5-5
…below its group

A group's consolidated metadata holds the hierarchy below the group:
zarr-python keeps the document of the node at /a/b of that hierarchy,
the group its root, at the key a/b. Nothing judged the keys or the tree
they make. A key that is no node's path ("", "__a", "a/../b", "/a"), a
document below an array's, and a document whose group is missing were
read without a problem; zarr-python raises a bare KeyError on three of
them and keeps the rest. Each is now a problem at its entry, a missing
group a missing_key, and ZarrV3ConsolidatedMetadata refuses them when it
is built. A listed group's own listing lists what the group lists, at
the joined key and of the same node type: a node it lists alone is
dropped by zarr-python's reader, which keeps the flat listing.

The rules are a hierarchy's, not consolidated metadata's: a private
hierarchy_problems judges the node type of each node, by its node path,
as the spec's tree, so a model of a whole hierarchy can use it too. Every
problem is one per value, saying its reasons at once -- a path of a
thousand bad names is one problem, a chain of missing groups one
missing_key at the nearest -- so what is reported weighs what was read.
NodeName and NodePath, in zarr_metadata.v3, are the spec's node names
and paths, modelled on zarrs' types, and zarr_metadata.model judges a
string by them with validate/is/parse_node_name_v3 and their node_path
twins.

Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-fable-5-1
…ment

A v3 array model compared its members with ==, so "NaN" and "0x7fc00000"
were two float32 fill values, 0.0 and -0.0 one, a blosc with and without
the typesize that noshuffle ignores two codecs though canonical_of spelled
them alike, and a model holding a NaN attribute never equalled its own
copy. Two models are now equal when they mean the same document: what
the package interprets compares by its canonical spelling, and what it
does not compares as JSON text, which tells true from 1 and -0.0 from
0.0 and takes NaN for itself.

A data type's definition says which spelling of a fill value is its
value's own, DataTypeDefinition.fill_value_canonical, and
canonical_fill_value spells one, or gives UNSET for a fill value with a
problem: the float types read a number as a float64, as JSON parsers and
numpy do, and spell a value by its shortest number, a named value, or a
NaN's bits; complex by part; the numpy time types' -2**63 as "NaT"; bytes
as base64; a struct field by field. A field, Read, Unclaimed or Refused,
compares by field_key: its definition and its canonical configuration as
text, with the fields it holds by their own keys; default, v2,
sharding_indexed and zstd gain a canonical folding their spec defaults.
Every model and field hashes as it compares.

Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-fable-5-1

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant