Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
52 changes: 39 additions & 13 deletions packages/zarr-metadata/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -71,7 +71,7 @@ against the chunk it is handed, a shard's inner and index codecs too.
The validators do no arithmetic on values: whether a fill value survives
a `cast_value` round trip is not judged.

Three choices the specs' words leave open, or settle two ways:
Two choices the specs' words leave open, or settle two ways:

- **`attributes` may hold `NaN`, `Infinity` and `-Infinity`.** The spec
interprets no attribute, and zarr-python and xarray write those numbers
Expand All @@ -85,11 +85,13 @@ Three choices the specs' words leave open, or settle two ways:
reads wrong bytes as surely as one that skips a data type reads wrong
values. It keeps its meaning on an unknown top-level member, which a
reader can skip.
- **A chunk length of 0 is allowed along a dimension of length 0.** The
core spec asks for non-zero chunk lengths only "when the corresponding
dimensions of the arrays have non-zero length"; the regular grid spec
says chunk sizes are greater than zero. The package follows the core
spec, which zarr-python 3.0 and 3.1 wrote for an empty dimension.

A regular grid's chunk lengths are at least 1, along a dimension of
length 0 too: "Chunk sizes must be greater than zero", the regular grid
spec says. The core spec's "non-zero when the corresponding dimensions
of the arrays have non-zero length" says less, and allows nothing more,
so a document with a 0 there, as zarr-python 3.0 and 3.1 wrote for an
empty dimension, is refused.

`read_array_metadata_v3` reads a document once and returns everything
the read found: each field as the scope read it -- `Read` by the
Expand Down Expand Up @@ -134,13 +136,37 @@ build the model of either kind, as the models' own `from_json` and
A member the spec does not define is not a field; the model's
`must_understand_fields` names those a reader must understand.

The Pydantic integration's generated JSON Schemas express independently
checkable document structure and field constraints, but they are not a
replacement for runtime model validation. Standard JSON Schema treats a
mathematically integral number such as `1.0` as an integer, while the runtime
boundary requires Python `int` values, and it cannot express arbitrary
same-length relations such as `dimension_names` versus `shape` or v2 `chunks`
versus `shape`. Consumers should run the model parser after schema validation.
`node_metadata_json_schema_v3` writes what the validators read as a
JSON Schema, draft 2020-12, for an editor that checks a `zarr.json` as it
is written, or a validator in another language. Each extension point is
a field as its scope reads it: a configuration as its definition's
TypedDict says, bounds and all, and a name nothing in the scope claims
with any configuration. The fill value is what the data type it names
takes. `field_json_schema(kind, context)`, in
`zarr_metadata.v3.definition`, writes one field's schema, and
`json_schema`, in `zarr_metadata.typed_json`, any TypedDict's, as `check`
reads it. A schema says what each member is, and not what the rules say
of members together, so a document it accepts may still have a problem;
a JSON document the validators accept, it accepts. A validator reads
JSON as a parser gives it, arrays as lists: a model's `to_json` writes
tuples, which a Python validator does not take for arrays.

```python
import json
from zarr_metadata.model import node_metadata_json_schema_v3

with open("zarr.schema.json", "w") as f:
json.dump(node_metadata_json_schema_v3(), f, indent=2)
```

The Pydantic integration's field types have JSON Schemas of their own,
for a model that holds them: an extension point there is a name and any
configuration, read in no scope, and v2 documents have one too. For a
`zarr.json`, use `node_metadata_json_schema_v3`. Neither replaces the
validators: JSON Schema takes a number such as `1.0` for an integer,
where the models require an `int`, and says nothing of what members read
together say, such as `dimension_names` against `shape` or v2 `chunks`
against `shape`. Run the model parser after schema validation.

## Scope

Expand Down
6 changes: 3 additions & 3 deletions packages/zarr-metadata/changes/371.doc.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,6 @@ The validation boundary says what the package decides where the specs
leave it open or settle it two ways: `attributes` may hold `NaN`,
`Infinity` and `-Infinity`, which `to_key_value` writes as bare tokens;
`must_understand: false` is refused at every extension point, codecs
too; a chunk length of 0 is allowed along a dimension of length 0. The
models' `from_json`, `to_json`, `from_key_value` and `to_key_value`
have docstrings, and no public docstring names a private function.
too. The models' `from_json`, `to_json`, `from_key_value` and
`to_key_value` have docstrings, and no public docstring names a private
function.
9 changes: 9 additions & 0 deletions packages/zarr-metadata/changes/378.bugfix.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
The `regular` chunk grid refuses a chunk length of 0, along a dimension
of length 0 too, as its specification says: "Chunk sizes must be greater
than zero". `RegularChunkGridConfiguration.chunk_shape` is
`tuple[Annotated[int, Ge(1)], ...]`, so a 0 is an `invalid_value` at its
place, with the bound in its `ctx`. The package had read the core
specification's "The chunk shape elements are non-zero when the
corresponding dimensions of the arrays have non-zero length" as allowing
0 on an empty dimension, as zarr-python 3.0 and 3.1 wrote it; that
sentence says less than the grid's own, not something else.
21 changes: 21 additions & 0 deletions packages/zarr-metadata/changes/378.feature.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
JSON Schemas of what the package reads, draft 2020-12, as pydantic's
`TypeAdapter(...).json_schema()` and zod's `toJSONSchema` write theirs.
`node_metadata_json_schema_v3(context=...)`, in `zarr_metadata.model`,
is a `zarr.json`'s, an array's or a group's, for an editor or a
validator in another language: each extension point a field as its
scope reads it -- a configuration as its definition's TypedDict says,
bounds and all, or a name nothing in the scope claims, with any
configuration -- the fill value what the data type it names takes, and
a group's consolidated metadata the documents it holds.
`field_json_schema(kind, context)`, in `zarr_metadata.v3.definition`, is
one field's, and `json_schema(shape)`, in `zarr_metadata.typed_json`,
any TypedDict's, as `check` reads it. What the rules say of members
together is not in a schema, so a document it accepts may still have a
problem; a JSON document the validators accept, it accepts. JSON Schema
takes `1.0` for an integer, where the package wants `1`. A v3 array
document's extension points are annotated with the field aliases --
`data_type: DataTypeField`, `codecs: tuple[CodecField, ...]` -- which
are the JSON a metadata field is, so a type checker and `check` read
them as before; its `shape` holds integers of at least 0, which `check`
now holds it to. A data type whose `fill_value` holds a metadata field
is refused when it is built: a fill value is a value of its data type.
9 changes: 4 additions & 5 deletions packages/zarr-metadata/changes/4443.feature.3.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,6 @@
**Breaking:** a v3 array's `chunk_grid` is judged against its `shape`, by
the grid's definition. A regular grid whose `chunk_shape` does not have
one length per dimension of the shape, or has a length of 0 for a
dimension that is not empty, and a rectilinear grid whose `chunk_shapes`
does not have one entry per dimension, or whose chunk lengths fall short
of their dimension, each have a problem in `chunk_grid.configuration`,
where the package accepted them before.
one length per dimension of the shape, and a rectilinear grid whose
`chunk_shapes` does not have one entry per dimension, or whose chunk
lengths fall short of their dimension, each have a problem in
`chunk_grid.configuration`, where the package accepted them before.
10 changes: 6 additions & 4 deletions packages/zarr-metadata/docs/api/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,12 +7,14 @@ title: API reference
The package is organized to mirror the structure of the Zarr specifications:

- [`zarr_metadata.model`](model.md) — frozen-dataclass document models,
validators, loc-aware parsers, and the `UNSET` sentinel
validators, loc-aware parsers, a `zarr.json`'s JSON Schema, and the
`UNSET` sentinel
- [`zarr_metadata.pydantic`](pydantic.md) — optional Pydantic field types
over the models
- [`zarr_metadata.typed_json`](typed_json.md) — `check`, which type-checks
a JSON value against any of the package's `TypedDict`s, read as the
typing spec defines them, with every problem located
typing spec defines them, with every problem located, and
`json_schema`, which writes what `check` reads as a JSON Schema
- [`zarr_metadata.v2`](v2.md) — `TypedDict` shapes for Zarr v2 documents
(`.zarray`, `.zgroup`, `.zattrs`, `.zmetadata`)
- [`zarr_metadata.v3`](v3/index.md) — `TypedDict` shapes for Zarr v3
Expand All @@ -22,8 +24,8 @@ The package is organized to mirror the structure of the Zarr specifications:
- [`zarr_metadata.v3.definition`](v3/definition.md) — each extension's
metadata as a definition: the TypedDict its configuration is, and the
rules on it; check JSON against a TypedDict, judge a configuration,
read a whole field in a scope, or read a codec pipeline. Its module
docstring is the guide
read a whole field in a scope, read a codec pipeline, or write a
scope's fields as a JSON Schema. Its module docstring is the guide

The document types, models, and spec vocabulary — including the store keys —
are re-exported at the top level, so
Expand Down
37 changes: 31 additions & 6 deletions packages/zarr-metadata/docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,7 +86,7 @@ against the chunk it is handed, a shard's inner and index codecs too.
The validators do no arithmetic on values: whether a fill value survives
a `cast_value` round trip is not judged.

Three choices the specs' words leave open, or settle two ways:
Two choices the specs' words leave open, or settle two ways:

- **`attributes` may hold `NaN`, `Infinity` and `-Infinity`.** The spec
interprets no attribute, and zarr-python and xarray write those numbers
Expand All @@ -100,11 +100,13 @@ Three choices the specs' words leave open, or settle two ways:
reads wrong bytes as surely as one that skips a data type reads wrong
values. It keeps its meaning on an unknown top-level member, which a
reader can skip.
- **A chunk length of 0 is allowed along a dimension of length 0.** The
core spec asks for non-zero chunk lengths only "when the corresponding
dimensions of the arrays have non-zero length"; the regular grid spec
says chunk sizes are greater than zero. The package follows the core
spec, which zarr-python 3.0 and 3.1 wrote for an empty dimension.

A regular grid's chunk lengths are at least 1, along a dimension of
length 0 too: "Chunk sizes must be greater than zero", the regular grid
spec says. The core spec's "non-zero when the corresponding dimensions
of the arrays have non-zero length" says less, and allows nothing more,
so a document with a 0 there, as zarr-python 3.0 and 3.1 wrote for an
empty dimension, is refused.

`read_array_metadata_v3` reads a document once and returns everything
the read found: each field as the scope read it -- `Read` by the
Expand Down Expand Up @@ -149,6 +151,29 @@ build the model of either kind, as the models' own `from_json` and
A member the spec does not define is not a field; the model's
`must_understand_fields` names those a reader must understand.

`node_metadata_json_schema_v3` writes what the validators read as a
JSON Schema, draft 2020-12, for an editor that checks a `zarr.json` as it
is written, or a validator in another language. Each extension point is
a field as its scope reads it: a configuration as its definition's
TypedDict says, bounds and all, and a name nothing in the scope claims
with any configuration. The fill value is what the data type it names
takes. `field_json_schema(kind, context)`, in
`zarr_metadata.v3.definition`, writes one field's schema, and
`json_schema`, in `zarr_metadata.typed_json`, any TypedDict's, as `check`
reads it. A schema says what each member is, and not what the rules say
of members together, so a document it accepts may still have a problem;
a JSON document the validators accept, it accepts. A validator reads
JSON as a parser gives it, arrays as lists: a model's `to_json` writes
tuples, which a Python validator does not take for arrays.

```python
import json
from zarr_metadata.model import node_metadata_json_schema_v3

with open("zarr.schema.json", "w") as f:
json.dump(node_metadata_json_schema_v3(), f, indent=2)
```

## Scope

At minimum, this library supports what Zarr-Python needs: the complete
Expand Down
Loading
Loading