Skip to content

feat(zarr-metadata)!: a shard is judged against the chunk it is handed, and its pipelines are read - #4443

Merged
d-v-b merged 17 commits into
zarr-developers:mainfrom
d-v-b:feat/zarr-metadata-sharding
Sep 28, 2026
Merged

d-v-b merged 17 commits into
zarr-developers:mainfrom
d-v-b:feat/zarr-metadata-sharding

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

🤖 AI text below 🤖

The last of four pieces of layer 4, depending on #4442 (4c, the codec pipeline): the sharding codec is judged against the shard it is handed, and its two pipelines are read as the array's is. The first fourteen commits are #4440's, #4441's and #4442's; this PR's are the last three.

document["shape"] = [16, 16]
document["chunk_grid"] = {"name": "regular", "configuration": {"chunk_shape": [8, 6]}}
document["codecs"] = [{"name": "sharding_indexed", "configuration": {
    "chunk_shape": [4, 4],
    "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}],
    "index_codecs": ["bytes"],
}}]
validate_array_metadata_v3(document)
# codecs.0.configuration.chunk_shape.1: expected an inner chunk length that divides each length the shards take along this axis, [6], got 4
# codecs.0.configuration.index_codecs.0.configuration.endian: expected an endian, since each 'uint64' value takes several bytes

The shard. The sharding codec's chunk_shape has an inner chunk length for each axis of the shard it is handed, and each length divides every length the shards take along that axis. The shard is the chunk the codec is handed, after any codec before it. A transposed shard is divided along its transposed axes, as decided for layer 4. zarr-python reads divisibility from the untransposed grid instead.

The pipelines it holds. A codec definition gains one more hook, pipelines(configuration, nested, chunk). For each member of its configuration that holds a pipeline, it gives the chunk that pipeline's first codec is handed:

  • the inner codecs are handed the inner chunks, of chunk_shape and of the shard's data type;
  • the index codecs are handed the shard index: uint64, with the chunks per shard along each axis and a final axis of 2.

Each pipeline is read as the array's is: its order, then each codec against the chunk it is handed, with problems where they sit. A Stage keeps the inner stages as inner. An index bytes codec without an endian is now caught, and so is an inner transpose of the wrong rank. This also holds behind a codec nothing in scope claims, since the index is uint64 whatever the shard.

Like transition, pipelines is asked whatever the chunk rules found, and gives only what holds. A shard of another number of axes than its chunk_shape hands both pipelines lengths nothing is known of, since which of the two is wrong is not known. So it draws one problem, not three.

A problem inside is the inner field's own. This is the first commit. resolve used to read a field as invalid when anything inside it had a problem. So a shard holding a gzip with a level out of range, or a codec with must_understand: false, was left unread, and nothing else about it was judged. A document's fields are not read that way: a bad codec at one index leaves the others read. Now the fields a configuration holds are read the same way:

  • a problem of one is its own, reported where it sits, with its own resolution in nested;
  • the field holding it is read when its own configuration and rules are sound;
  • its rules are still asked only when each field it holds is named, since a rule may read one by name.

Every hook already treats an unread nested field as unknown, so this adds no false problems. It changes what resolve gives for a container (layer 2, merged, unreleased); two tests encoded the old answer. Without it, the correctness review found problems missing from 553 of 3,000 generated documents with shards, although the verdicts were right. The same codec also drew different problems at the top level and inside a shard.

Also here. The codec functions no codec of a kind is asked are now one declared table, _UNASKED, instead of one if each. A bytes -> bytes codec is handed bytes, and only an array -> array codec hands on a chunk.

Breaking. A document whose shard does not divide, or whose shard pipelines do not fit what they are handed, now has a problem where it sits, where the package accepted it before. A field holding a field with a problem now resolves as read, where it resolved as invalid.

Choices worth a look

  • The index's data type is a hand-built reading of the package's uint64 definition. Hooks get no scope, and the spec fixes the index type whatever the scope holds.
  • Pipelines are keyed by the member of the configuration that holds them. No codec holds a pipeline deeper than that.
  • Handed a chunk nothing is known of, a shard still hands on the inner chunks it declares, and an index of len(chunk_shape) + 1 axes. That is its own declaration, as cast_value names its data type.
  • A shard's pipelines are judged only as a document's codecs are read. resolve or canonicalize of a shard field alone does not see that codecs: [] has no array -> bytes codec. That rule needs no chunk, so it could move into the shard's rules, with the walk skipping it for held pipelines. Left for a follow-up.
  • A pipelines naming a member that holds data type fields nothing in scope claims cannot be told from one holding codecs nothing claims. It reads them as codecs nothing claims. Only a custom definition can do this.

Reviews. Two, with separate lenses.

The correctness review extended the 4c spec reference with the sharding rules:

  • 120,000 generated documents, with every inner and index stage compared, gave 0 disagreements against a reference that mirrors the walk.
  • All 17 planted bugs were caught, two only by the stage comparison, since no rule reads the index lengths.
  • 33,000 fuzzed documents raised nothing.
  • A strict reference found the problems hidden inside shards, which the first commit fixes. It also found a rank mismatch that drew three problems, now one.
  • It found a divisibility message that listed every length: on a rectilinear axis of 20,000 edges it was 134,548 characters, and it now names the smallest length not divided.

The design review's changes are in:

  • pipelines are keyed by member, not by path, so an untested path walk is gone;
  • a member holding no codecs is a TypeError;
  • the rank mismatch is fixed;
  • the tests are split per case, and the nested shard is folded into the main table.

Where zarr-python differs. The spec is the authority. The first two are filed: the transpose case on #2050, and the shard inside a shard as #4437.

  • zarr-python judges divisibility on the untransposed grid. It accepts [transpose(1,0), sharding_indexed chunk_shape [4,3]] on chunks of (4, 6), and a write and read returns different data. It refuses (6, 4), which its runtime handles.
  • It never checks a shard held inside a shard. It accepts a shard [4,4] holding a shard [3,4], and a round trip returns different data.
  • It accepts index_codecs of [crc32c], of [bytes, bytes], of bytes without an endian, or with a transpose of the wrong rank, and fails at the first write.

The stack. Layers 0 (#4420, #4421, #4422), 1 (#4432), 2 (#4434) and 3 (#4436) have merged. 4a is #4440, 4b is #4441, 4c is #4442, and this is 4d, which completes layer 4.

🤖 Generated with Claude Code

Each data type's definition declares the JSON shape of its fill value and
the rules for one of that shape, and the v3 array validators judge a
document's fill value against the data type the scope read. A field that
is read keeps the fields it read inside, by location, so a struct judges
each field's fill value by that field's own type.

Assisted-by: ClaudeCode:claude-opus-5-5
…resolve kept

The simplest spelling of a field that read spells each field it holds
from its reading in `Resolved.nested`, rather than reading every subtree
again at each level. A nested field of a field without problems has none,
so the branch for one without a spelling could not be taken.

Assisted-by: ClaudeCode:claude-opus-5-5
…see past unknown keys

- The JSON check walks one frame per level of nesting, as `refine_json`
  does, so a fill value checked against a JSON-typed shape reads as deep
  as `refine_json` reads: a struct fill value 600 levels deep raised
  `RecursionError`.
- The integer, float and complex fill value rules are partial applications
  of module-level functions, so a scope's definitions pickle again.
- A key a fill value's shape does not declare is reported and left out, and
  the fill value rules still judge the rest, as `judge` does for a
  configuration.
- A struct whose reading holds no reading of a field's type leaves that
  field's fill value unjudged.
- `fill_value_problems` takes the reading of any data type, whatever its
  configuration's type.
- `create_default` says that an overridden `data_type` needs its own
  `fill_value`.

Assisted-by: ClaudeCode:claude-opus-5-5
A chunk grid's definition says which arrays it fits, and the v3 array
validators judge a document's grid against its shape once both are read:
a regular grid has a chunk length for each dimension, and 0 only for a
dimension of length 0; a rectilinear grid has chunk lengths for each
dimension that cover it.

Assisted-by: ClaudeCode:claude-opus-5-5
…ength of 1

The grid `create_default` derives from an overridden `shape` had a chunk
length of 0 for a dimension of length 0. The regular grid asks for chunk
lengths greater than zero, and `zarr` refuses to open such a grid, so the
derived grid now has a chunk length of 1 there. It still fits the shape.

Assisted-by: ClaudeCode:claude-opus-5-5
Assisted-by: ClaudeCode:claude-opus-5-5
…ration holds

`rules` is handed the configuration and the fields it holds as the scope
read them (`Nested`), as `fill_value_rules` and `shape_rules` are, so a
rule can judge a configuration by what is inside it: a struct's field
types, a cast's target. `resolve` reads the nested fields before it asks
the rules, and still reports the rules' problems first; `judge` reads in
no scope, and hands them nothing read. Every definition's rules take the
new argument.

Assisted-by: ClaudeCode:claude-opus-5-5
… the chunk it is handed

A v3 array's codecs are read in order, and each codec handed an array is
judged against the chunk it is handed. The array hands its first codec a
`Chunk`: the lengths its grid's chunks take along each of the array's
dimensions, which the grid's new `chunk_lengths` says, and its data
type. A codec definition says what the spec disallows in it handed a
chunk, `chunk_rules`, and an array -> array codec says what it hands on,
`transition`, whatever its chunk rules found. `transpose` has both: its
`order` has one entry per axis of its chunk, and it hands on the axes
permuted.

`read_pipeline(codecs, chunk)` judges the order first (array -> array
codecs, one array -> bytes codec, bytes -> bytes codecs), which refuses
an empty list too, and gives each codec's `Stage` with the chunk it is
handed. Nothing is guessed: the codec after one the scope did not read,
or after one that says nothing of what it hands on, is handed a chunk
nothing is known of. `chunk_grid_lengths` replaces the private
`chunk_grid_problems`, giving a grid field's lengths with its problems,
and a `Chunk` checks what it holds, so a definition's fault is reported
as the definition's.

Assisted-by: ClaudeCode:claude-opus-5-5
A data type definition says how its values are stored, `storage`: in
single bytes, in several bytes at a time, or each in as many as it
needs. `bool`, `int8`, `uint8` and `r8` are made of single bytes; the
other core types, and the numpy time types, hold numbers of several
bytes; `string` and `bytes` vary. Of raw bits wider than a byte the spec
does not say whether a byte order applies, so their storage is unknown.
A struct's is its fields', packed together. `storage_of` asks it of a
data type field a scope read.

Two rules follow. The `bytes` codec takes an `endian` when it is handed
numbers of several bytes, and is not handed values that vary in size. A
struct's field types are of fixed size.

Assisted-by: ClaudeCode:claude-opus-5-5
…the data types they meet

`cast_value` casts to a data type that models real numbers, wraps only to
an integral one, and is handed one that models real numbers; each scalar
of its `scalar_map` is a fill value of the data type on its side of the
cast. It hands on its data type, whatever it is handed. `scale_offset`
is handed a data type with arithmetic, and its `offset` and `scale` are
fill values of it; it hands on what it is handed.

Each codec names the data types it takes "defined in this repository",
so one it does not name is refused only where the spec's own words refuse
it: truth values, raw bits, text, byte strings and records are no
numbers, and complex numbers model no real number. The numpy time types
say they take any codec of 64-bit integers, and scale_offset does not
say whether complex numbers are its, so those are left be. A scalar is
judged by the core fill value encoding, which writes a float's infinity
`"Infinity"`, where the cast_value codec's own example writes
`"+Infinity"`.

Assisted-by: ClaudeCode:claude-opus-5-5
…till claimed by its name

`resolve` read a field whose configuration was not an object as claimed
by nothing, although its name names a definition, so the codec order was
judged without it: `[{"name": "gzip", "configuration": 5}, "bytes"]` drew
no problem for the gzip before the bytes codec. The field keeps the
definition its name names, unread.

Assisted-by: ClaudeCode:claude-opus-5-5
…s own

`resolve` read a field as invalid when anything inside it had a problem:
a codec a shard holds with a gzip `level` out of range, or a
`must_understand` of false on one, left the shard unread, so nothing
else about it was judged. A document's fields are not read that way: a
problem with one codec leaves the others read. Now the fields a
configuration holds are read the same way. A problem of one is its own,
reported where it sits with its own resolution in `nested`, and the field
holding it is read when its own configuration and rules are sound. Its
rules are still asked only when each field it holds is named, since a
rule may read one by name.

Assisted-by: ClaudeCode:claude-opus-5-5
…d, and its pipelines are read

The sharding codec's `chunk_shape` has an inner chunk length for each
axis of the shard it is handed, dividing each length the shards take
along it. The shard is the chunk the codec is handed, so a transposed
shard is divided along its transposed axes.

A codec definition says what the pipelines it holds are handed,
`pipelines`, by the member of its configuration that holds each: the
shard's inner codecs are handed its inner chunks, of its data type, and
its index codecs the shard index, of `uint64` with an axis more than the
shard. Each is read as the array's pipeline is, its problems where it
sits, and a `Stage` keeps its stages as `inner`. Like the transition,
`pipelines` gives only what holds whatever the chunk rules found: a
shard of another number of axes than its `chunk_shape` hands both
pipelines lengths nothing is known of. The functions no codec of a kind
is asked are one declared table.

Assisted-by: ClaudeCode:claude-opus-5-5
@read-the-docs-community

Copy link
Copy Markdown

Documentation build overview

📚 zarr-metadata | 🛠️ Build #34798189 | 📁 Comparing 074c7c9 against latest (0401a7f)

  🔍 Preview build  

7 files changed · ± 7 modified

± Modified

@codecov

codecov Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.37%. Comparing base (0401a7f) to head (074c7c9).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #4443   +/-   ##
=======================================
  Coverage   94.37%   94.37%           
=======================================
  Files          93       93           
  Lines       13174    13174           
=======================================
  Hits        12433    12433           
  Misses        741      741           
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@d-v-b
d-v-b marked this pull request as ready for review September 28, 2026 08:31
@d-v-b
d-v-b merged commit 47528fe into zarr-developers:main Sep 28, 2026
40 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs release notes Automatically applied to PRs which haven't added release notes zarr-metadata Specific to the zarr-metadata sub-package

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant