Skip to content

feat(zarr-metadata)!: export JSON Schemas of a TypedDict, a field in a scope, and a zarr.json - #378

Open
d-v-b wants to merge 2 commits into
feat/zarr-metadata-field-problemsfrom
feat/zarr-metadata-json-schema
Open

d-v-b wants to merge 2 commits into
feat/zarr-metadata-field-problemsfrom
feat/zarr-metadata-json-schema

Conversation

@d-v-b

@d-v-b d-v-b commented Sep 29, 2026

Copy link
Copy Markdown
Owner

🤖 AI text below 🤖

JSON Schemas of what the package reads, draft 2020-12, as pydantic's
TypeAdapter(...).json_schema() and zod's toJSONSchema write theirs.
node_metadata_json_schema_v3(context=...), in zarr_metadata.model,
is a zarr.json's, an array's or a group's, for an editor or a
validator in another language: each extension point a field as its
scope reads it -- a configuration as its definition's TypedDict says,
bounds and all, or a name nothing in the scope claims, with any
configuration -- the fill value what the data type it names takes, and
a group's consolidated metadata the documents it holds.
field_json_schema(kind, context), in zarr_metadata.v3.definition, is
one field's, and json_schema(shape), in zarr_metadata.typed_json,
any TypedDict's, as check reads it. What the rules say of members
together is not in a schema, so a document it accepts may still have a
problem; a JSON document the validators accept, it accepts. JSON Schema
takes 1.0 for an integer, where the package wants 1. A v3 array
document's extension points are annotated with the field aliases --
data_type: DataTypeField, codecs: tuple[CodecField, ...] -- which
are the JSON a metadata field is, so a type checker and check read
them as before; its shape holds integers of at least 0, which check
now holds it to. A data type whose fill_value holds a metadata field
is refused when it is built: a fill value is a value of its data type.

The regular chunk grid refuses a chunk length of 0, along a dimension
of length 0 too, as its specification says: "Chunk sizes must be greater
than zero". RegularChunkGridConfiguration.chunk_shape is
tuple[Annotated[int, Ge(1)], ...], so a 0 is an invalid_value at its
place, with the bound in its ctx. The package had read the core
specification's "The chunk shape elements are non-zero when the
corresponding dimensions of the arrays have non-zero length" as allowing
0 on an empty dimension, as zarr-python 3.0 and 3.1 wrote it; that
sentence says less than the grid's own, not something else.

Three functions, each beside the reader it describes

function writes beside
json_schema(shape), zarr_metadata.typed_json a TypedDict, as check reads it check
field_json_schema(kind, context), zarr_metadata.v3.definition one field of a kind, as a scope reads it resolve
node_metadata_json_schema_v3(context=...), zarr_metadata.model a zarr.json validate_node_metadata_v3

What a schema holds

  • A field in a scope is one of three things:

    • each definition's field: its name as a const, its configuration's TypedDict in $defs, must_understand: true if present, and the bare name when no configuration is required;
    • raw bits: r and digits, matched to the end of the name, with nothing configured beside it;
    • a name none of the definitions is written with, bare or with any configuration.

    A shard's index_codecs takes codecs of static size, and names nothing claims.

  • A zarr.json holds its fill value to the data type it names, with one if/then per data type in scope: "int8" makes the fill value an integer in [-128, 127]. A group's consolidated_metadata holds array and group documents, or is null.

  • Not in any schema: the rules. Those include one dimension name per dimension, a grid that fits the shape, the pipeline's order, each codec against its chunk, blosc's typesize against shuffle, and a hex fill value's width. So a schema is sound rather than exact.

  • Layout. Each TypedDict and type alias goes in $defs under its name (GzipCodecConfiguration, CodecField, ZarrV3ArrayMetadataJSON), numbered on a clash. The root is written in place unless something refers to it. No part of a returned schema is shared with another.

Changes that come with it

  • Field aliases. The six field aliases move to v3._common, and ZarrV3ArrayMetadataJSON and its Partial annotate their extension points with them. The model's tables of extension points are now read off that TypedDict, so one declaration says which member is which kind.
  • shape is tuple[Annotated[int, Ge(0)], ...].
  • Fill values. A data type whose fill_value holds a field alias is refused when built. The checker reads a fill value as a value, and a schema would write it as a field.

What changes for a caller

before now
a regular grid's chunk_shape: [0, 4] over shape: [0, 4] valid invalid_value at chunk_shape.0, ctx {"ge": 1}
a 0 over a dimension of length 128 expected a chunk length >= 1 for a dimension of length 128, got 0 expected an integer >= 1, got 0
check(doc, ZarrV3ArrayMetadataJSON) with a negative dimension no problem invalid_value at ("shape", i)
typeddict_keys(ZarrV3ArrayMetadataJSON).members["data_type"] str | ZarrV3NamedConfigJSON DataTypeField, an alias of it
DataTypeDefinition(..., fill_value=tuple[CodecField, ...]) built TypeError

Spec correction

The regular-grid commit follows a correction at review. The agent had justified a test change by saying a regular grid takes a chunk length of 0 on an empty dimension. The grid's own page says otherwise: "Chunk sizes must be greater than zero" (chunk-grids/regular-grid/index.rst L40, since zarr-specs f538382). The README had listed following the core sentence as a choice the package made; that choice is gone. create_default already wrote 1 there.

Verification

  • Tests. 1,657 pass on Python 3.14, and 1,655 plus 2 skipped on 3.11. pyright, ruff and the docs build are clean.
  • Properties.
    • The typed_json schema agrees with the checker's reference over the harness's drawn types and values, reading 1.0 as 1: 300 examples in CI, 10,000 locally.
    • Bounds layered through Annotated, NewType and type aliases are written as check holds a value to them.
    • Fields and documents changed in one or two places: whatever the package reads without a problem, the schema accepts.
  • Soundness over the review corpus. Each document and field was round-tripped through json, then judged in core, core-and-extensions and an empty scope.
    • 587,487 and 759,096 judgments: in none does the package accept while the schema refuses.
    • The schema refuses 79–82% of what the package refuses. The rest is the rules', led by blosc typesize (6,564 fields), a grid that doesn't fit its shape (1,267 documents), pipeline order, index codec endian, fill value rules, and zarr_format: 3.0, which JSON Schema takes for 3.
  • Stack depth. Two limits turned up, neither of them a schema verdict:
  • Cost. Writing the node schema takes 0.7 ms and gives 24 KB, with 28 $defs. jsonschema validates a typical document in 0.56 ms; validate_node_metadata_v3 takes 0.04 ms.

The zarr-extensions registry, cross-checked

The copy is the one vscode-zarr vendors, at 4da7b37a, checked over the corpus fields for the 32 names both define.

  • The registry refuses what the package reads:
    • must_understand, at every name;
    • the bare name of every field that takes one: "int8", "crc32c", "default", "v2", "scale_offset";
    • {"name": "int8", "configuration": {}};
    • a shard's empty inner codecs.
  • The registry accepts what the types refuse:
    • a regular grid's 0;
    • unknown configuration keys and stray envelope members in rectilinear;
    • must_understand: false;
    • problems inside nested fields, which it doesn't judge: a shard's gzip level: 99, a dynamic index codec, a struct field's scale_factor: 0.
  • Everything else is a rule neither schema says.

Nothing was filed upstream.

Scenario re-runs

The recordings are updated in the untracked copy on this branch's worktree.

  • OME-Zarr library. 7 of 7 blocks unchanged.
  • zarr-python swap. The two readers agree on 136 of 158 documents, up from 131. The five new agreements are the documents with a 0 on an empty dimension, as zarr 3.0.10 and 3.1.6 wrote them, which both now refuse.
  • Metadata cleaner. Same repaired files. Four blocks show a zero chunk length's message as the bound's.
  • New codecs. valid_reshape_zero_length_dimension.json now chunks its empty dimension by 1, as the spec asks, and reads again.

Reviews

  • roborev, first pass. It found:
    • a clause in 4443.feature.3.md that described the old zero-length rule (dropped);
    • scenario recordings that still read a 0 (updated);
    • no test of check holding a document's shape to its bound (added).
  • roborev, second pass. Schemas.defined left a reserved entry behind when writing it failed. It now gives the reservation up; a test pins it.
  • Design review. Correctness findings:
    • The raw-bits pattern ended in $, which a Python validator also matches before a final newline. So {"name": "r16\n", "configuration": {"x": 1}}, which names nothing, was refused. It now ends in (?![\s\S]), and table rows pin it; the mutation tests rarely reach that case on their own.
    • The claim "a document the validators accept, the schema accepts" was false for to_json()'s tuples. It now says JSON as a parser gives it.
    • A fill value holding a field alias was read as JSON but written as a field. It is now refused.
    • Bounds on a bounded NewType were merged by overwriting. They now take the stricter, with a property test.
  • Design review, API shape. Its findings led to:
    • the extension-point tables being derived from the TypedDict;
    • the unclaimed branch written from written_name, one declaration of how a document names a definition;
    • returned schemas sharing nothing;
    • a typing-only isinstance becoming a cast;
    • the README saying which JSON Schema to use.

Choices worth a look

  1. Sound, not exact. Schemas leave the rules out rather than approximating them.
  2. Descriptions come from a Doc only, as in zod, not from docstrings, as in pydantic, because the package's docstrings are written for Python readers. Nothing in the package carries a Doc yet, so its own schemas have no descriptions.
  3. One document function. An array's or a group's schema alone is $defs/ZarrV3ArrayMetadataJSON or $defs/ZarrV3GroupMetadataJSON.
  4. The grid fix is in this PR as its own commit. It refuses documents zarr 3.0 and 3.1 wrote, which zarr-python itself already refuses.
  5. JSONSchema joins the public naming grammar's standalone vocabulary, beside JSONValue.

Left out

  • consolidated_metadata on ZarrV3GroupMetadataJSON. The spec now names this member. Declaring it would remove _group and additional_reserved_keys, but needs the consolidated TypedDict to move into group.py, and changes a public document type.
  • The pydantic types emitting this schema. They keep the scope-free one from _pydantic_schema.py.
  • A committed schema artifact for zarr-metadata-ts or vscode-zarr. vscode-zarr's generate_schemas.py reads the private _pydantic_schema today.
  • Smaller gaps. v2 schemas; the pipeline's order by contains/maxContains; Docs on the document TypedDicts.
Stack

Stacked on #377 (feat/zarr-metadata-field-problems), itself on #376 → #374 → #373 → #372 → #371 → #370. Merge after #377. Fragments: 378.feature.md and 378.bugfix.md; this PR also drops the now-untrue zero-length clauses from 371.doc.md and 4443.feature.3.md. Two commits: the grid fix first, then the schemas.

🤖 Generated with Claude Code

…s its specification says

"Chunk sizes must be greater than zero"
(https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-grids/regular-grid/index.rst#L40),
along a dimension of length 0 too. `RegularChunkGridConfiguration`
bounded its lengths with `Ge(0)`, and its shape rules took a 0 on an
empty dimension, reading the core specification's "The chunk shape
elements are non-zero when the corresponding dimensions of the arrays
have non-zero length" as allowing it. That sentence says less than the
grid's own. The bound is now `Ge(1)`, so a 0 is an `invalid_value` at
its place wherever it is, and the shape rules check only that the grid
has the array's dimensionality. `create_default` already wrote 1 there.

BREAKING CHANGE: a regular grid with a chunk length of 0 on a dimension
of length 0, as zarr-python 3.0 and 3.1 wrote one, is now refused.

Assisted-by: ClaudeCode:claude-opus-5-5
… and a zarr.json

As pydantic's `TypeAdapter(...).json_schema()` and zod's `toJSONSchema`
write theirs, in draft 2020-12:

- `json_schema(shape)`, in `zarr_metadata.typed_json`, writes a
  TypedDict as `check` reads it. A bound is JSON Schema's keyword for
  it, the stricter where a type and its `NewType` both say one, and a
  `Doc` the `description`. Each TypedDict and type alias is written
  once, in `$defs`, under its name.
- `field_json_schema(kind, context)`, in `zarr_metadata.v3.definition`,
  writes a field of one kind as a scope reads it: each definition's
  field, and a name none of them is written with, with any
  configuration. Raw bits' name is matched to its end, so a validator
  that matches patterns as Python does takes no final newline for it.
- `node_metadata_json_schema_v3(context=...)`, in `zarr_metadata.model`,
  writes a `zarr.json`: its fill value held to the data type it names,
  and a group's consolidated metadata holding documents.

`Schemas` writes each shape, asking a caller's `SchemaLeaf` first, as
the checker asks a `Leaf`, and hands back a schema sharing nothing. A
schema says what the types say and not what the rules say, so every
JSON document the package finds nothing wrong with, the schema accepts.
Property tests hold the typed_json schema to the checker's reference,
bounds in layers to what `check` holds a value to, and fields and
documents changed in one or two places to soundness. JSON Schema takes
`1.0` for an integer. Only a `Doc` is a description, as in zod: the
package's docstrings are written for Python's readers.

The six field aliases move to `v3._common`, so `ZarrV3ArrayMetadataJSON`
says what each extension point is: `data_type: DataTypeField`,
`codecs: tuple[CodecField, ...]`. To a type checker and to `check` they
are the JSON they were, and the model's tables of extension points are
read off the TypedDict. `shape` holds integers of at least 0, which
`check` now holds it to. A data type whose `fill_value` holds a metadata
field is refused when it is built, since the checker reads a fill value
as a value and a schema would write it as a field.

Assisted-by: ClaudeCode:claude-opus-5-5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant