diff --git a/.github/workflows/zarr-metadata.yml b/.github/workflows/zarr-metadata.yml index 16b19181a7..f38280f25d 100644 --- a/.github/workflows/zarr-metadata.yml +++ b/.github/workflows/zarr-metadata.yml @@ -102,7 +102,7 @@ jobs: repo: casey/just version: 1.58.0 - name: Run pyright - # The pyright version and interpreter pins live in the justfile. + # Pyright runs unpinned; the interpreter pin lives in the justfile. run: just typecheck docs: diff --git a/packages/zarr-metadata/README.md b/packages/zarr-metadata/README.md index 1febfe80c3..db48871bf3 100644 --- a/packages/zarr-metadata/README.md +++ b/packages/zarr-metadata/README.md @@ -13,9 +13,9 @@ Two layers and an optional integration: and [Zarr v3](https://zarr-specs.readthedocs.io/en/latest/v3/core/index.html) specifications, plus types for [`zarr-extensions`](https://github.com/zarr-developers/zarr-extensions/) and a few widely-used-but-unspecified entities (e.g. consolidated metadata). -- **Document models** (`zarr_metadata.model`): canonical frozen-dataclass - models of whole metadata documents, with validators, loc-aware - parsers, and store-key (de)serialization. A document produced by `to_json` +- **Document models** (`zarr_metadata.model`): each model is a metadata + document and the scope it was read in, with validators, loc-aware + readers, and store-key (de)serialization. A document produced by `to_json` shares no mutable state with the model that produced it. - **Optional Pydantic integration** (`zarr_metadata.pydantic`, requires Pydantic 2.13 or newer): each model as a Pydantic field type that validates @@ -44,7 +44,8 @@ parser and returns the same normalized model class: from pydantic import TypeAdapter import zarr_metadata.pydantic as zmp -metadata = TypeAdapter(zmp.ZarrV3ArrayMetadata).validate_python(raw) +adapter: TypeAdapter[zmp.ZarrV3ArrayMetadata] = TypeAdapter(zmp.ZarrV3ArrayMetadata) +metadata = adapter.validate_python(raw) encoded = metadata.to_key_value()["zarr.json"] ``` @@ -52,12 +53,25 @@ A bare `TypeAdapter` over a public document `TypedDict` is a coercive shape adapter, not a Zarr conformance validator; it may coerce values or discard members that the strict model parser rejects. +## Dependencies + +The core of this package depends on `typing-extensions` and +`annotated-types` only. It does not depend on a validation framework, +and it will not: a dependency on pydantic, or any other framework with +its own release cadence and compiled parts, would pin that framework for +every consumer of zarr-metadata and collide with the pins consumers +already carry. Instead the package reads a `TypedDict` as the typing spec +defines it with its own checker (`zarr_metadata.typed_json.check`), and +spells bounds in the `annotated-types` vocabulary, which pydantic reads +too. `zarr_metadata.pydantic` is an optional integration over the models; +nothing in the core imports it. + ## Validation boundary The model validators enforce the declared document structure and a small set of context-free consistency rules, including fixed format literals, finite -JSON numbers, non-negative dimensions, and one `dimension_names` entry per -array dimension. In a v3 document they also read +JSON numbers outside `attributes`, non-negative dimensions, and one +`dimension_names` entry per array dimension. In a v3 document they also read each extension point -- the data type, chunk grid, chunk key encoding, each codec and each storage transformer -- through the definition that claims its name in a scope, `CORE_AND_EXTENSIONS` unless a `context` is passed: a @@ -71,13 +85,163 @@ against the chunk it is handed, a shard's inner and index codecs too. The validators do no arithmetic on values: whether a fill value survives a `cast_value` round trip is not judged. -The Pydantic integration's generated JSON Schemas express independently -checkable document structure and field constraints, but they are not a -replacement for runtime model validation. Standard JSON Schema treats a -mathematically integral number such as `1.0` as an integer, while the runtime -boundary requires Python `int` values, and it cannot express arbitrary -same-length relations such as `dimension_names` versus `shape` or v2 `chunks` -versus `shape`. Consumers should run the model parser after schema validation. +Two choices the specs' words leave open, or settle two ways: + +- **`attributes` may hold `NaN`, `Infinity` and `-Infinity`.** The spec + interprets no attribute, and zarr-python and xarray write those numbers + there (a CF `_FillValue`, say). The models read them, and `to_key_value` + writes them back as those bare tokens, which a strict JSON parser + refuses. `check`, from `zarr_metadata.typed_json`, refuses a non-finite + number wherever it is, attributes included. +- **A reader walks 256 levels of nesting.** A value nested deeper is an + `invalid_value` at the level past the last, wherever it sits. Every + reader, writer and comparison takes one frame for each level, and + `copy.deepcopy`, and `pickle` before Python 3.12, two: a document at + the cap takes about half of the interpreter's default limit, and the + rest is the caller's. +- **`consolidated_metadata: null` is a problem.** zarr-python 3.0 and 3.1 + wrote it on a group they had not consolidated; the spec says an object, + and the package models nothing else as right. + `read_repaired_node_metadata_v3` removes it before reading. +- **An extension is named as the spec names one**, `^[a-z][a-z0-9-_.]+$`, + or by a URI, which earlier versions of the spec required; any other + name is refused before a definition is asked, so `""` and `"foo/bar"` + are not unknown extensions but problems. +- **`must_understand: false` is refused at every extension point**, codecs + and storage transformers too, though the core spec names only the data + type, chunk grid and chunk key encoding: a reader that skips a codec + reads wrong bytes as surely as one that skips a data type reads wrong + values. It keeps its meaning on an unknown top-level member, which a + reader can skip. + +A regular grid's chunk lengths are at least 1, along a dimension of +length 0 too: "Chunk sizes must be greater than zero", the regular grid +spec says. The core spec's "non-zero when the corresponding dimensions +of the arrays have non-zero length" says less, and allows nothing more, +so a document with a 0 there, as zarr-python 3.0 and 3.1 wrote for an +empty dimension, is refused. + +`read_array_metadata_v3` reads a document once and returns everything +the read found: each field as the scope read it -- `AcceptedField` by the +definition that claims its name, `UnclaimedField` when none does, or +`RefusedField` -- with where it sits and the kind it was read as, each codec +with the chunk it is handed, every problem, and the model when there is +none; `from_json` is that model, or the problems raised. A consumer's +own policy is a walk over the fields, with nothing read twice: +`with_problems` gives each with its problems, those located in it and in +the fields it holds, as zod's `flattenError` groups issues, and +`canonical_of` spells a field with none in the fewest words, without +reading it again. Which fields go beyond the core spec, say -- a field +that names nothing is a problem already -- and how each is spelled most +simply: + +```python +from zarr_metadata.model import read_array_metadata_v3 +from zarr_metadata.v3.definition import CORE, canonical_of, with_problems + +reading = read_array_metadata_v3(raw) +beyond_core = [ + loc + for loc, field in reading.fields() + if field.name is not None and CORE.claimant(field.read_as, field.name) is None +] +simplest = { + loc: canonical_of(field, problems) + for loc, field, problems in with_problems(reading.fields(), reading.problems) +} +metadata = reading.metadata # None when reading.problems is not empty +``` + +`read_group_metadata_v3` reads a group the same way, and each document +its consolidated metadata holds once. `read_node_metadata_v3` reads a +`zarr.json` of either kind as the node its `node_type` says it is, as +a discriminated union reads its tag: a document that says neither reads +as `ZarrV3UnknownNodeReading`, with the problems, and nothing else of it +is read but its `zarr_format`, so a document of another format says it +is not v3. `node_metadata_from_json_v3` and `node_metadata_from_key_value_v3` +build the model of either kind, as the models' own `from_json` and +`from_key_value` build one. + +Consolidated metadata holds the hierarchy below its group, the group its +root: the document of the node at `/a/b` sits at the key `a/b`, and the +documents and the group make a tree in which only groups hold nodes and +each node's parent is held. `NodeName` and `NodePath`, in +`zarr_metadata.v3`, are the strings the spec's rules for node names and +paths hold of, modeled on zarrs' types of those names, and +`validate_node_name_v3`, `is_node_name_v3` and `parse_node_name_v3`, and +their `node_path` twins, judge a string by them. + +A member the spec does not define is not a field; the model's +`must_understand_fields` names those a reader must understand. + +A v3 model is its document and the scope it was read in: +`ZarrV3ArrayMetadata(document, context=None)` reads the document in the +scope, `CORE_AND_EXTENSIONS` when none is given, and raises +`MetadataValidationError` with every problem, so no model is built +invalid; `to_json` writes the document as it was written, and +`to_key_value` writes it as it is. Every typed member is a view of that +read: each field as the scope read it, an `AcceptedField` or an `UnclaimedField`, and +`shape`, `attributes` and the rest as the read refined them, read-only +at every level: a list given for an array as a tuple, an object as a +read-only mapping; `to_json` gives plain containers. A model is changed +by `update`, which puts +JSON members in place of the document's and reads the result in the +model's own scope, so no scope is passed back in; `with_context` reads +the document in another scope, and `refined_in` only in one that claims +what this one left unclaimed and contradicts nothing, raising +`ScopeConflictError` otherwise. The documents a group's +`consolidated_metadata` holds are models of the group's scope, built from +the group's one read; it takes node models as entries too, each accepted +when its claims refine into the group's scope and refused at its path +otherwise, and `Context.joined` is the scope to consolidate children of +several scopes in. Every reader takes `context=None` for the default +scope, `CORE_AND_EXTENSIONS`. + +Two models are equal when they mean the same document, however each is +spelled. What the package interprets -- each field, and the fill value +against its data type -- compares by its canonical spelling, as +`canonical_of` and `canonical_fill_value` give it: `"NaN"` and +`"0x7fc00000"` are one `float32` fill value, `0.0` and `-0.0` two, and a +blosc with and without the `typesize` that `noshuffle` ignores one +codec, and a v2 `dtype` by its family and size, `=0.6", # >=4.16: first release where `Sentinel` pickles by reference # (`__reduce__` returns the sentinel's name), so `UNSET` — and any model # holding it — can cross process boundaries and be deep-copied with its @@ -135,8 +139,7 @@ checks = [ "PR06", ] -# The pyright version lives in the justfile, which is what CI runs; pinning it -# keeps a pyright release from turning CI red on its own schedule. +# Pyright runs unpinned, from the justfile, which is what CI runs. [tool.pyright] include = ["src", "tests"] # Pyright is the only checker that runs on this package (the root mypy @@ -212,3 +215,7 @@ showcontent = true directory = "misc" name = "Misc" showcontent = true + +[tool.ruff.lint.flake8-type-checking] +# A TypedDict's annotations are evaluated at run time: the checker reads them. +runtime-evaluated-base-classes = ["typing_extensions.TypedDict", "typing.TypedDict"] diff --git a/packages/zarr-metadata/src/zarr_metadata/__init__.py b/packages/zarr-metadata/src/zarr_metadata/__init__.py index 1a6b39f04d..a227609c19 100644 --- a/packages/zarr-metadata/src/zarr_metadata/__init__.py +++ b/packages/zarr-metadata/src/zarr_metadata/__init__.py @@ -14,23 +14,23 @@ ProblemKind, ValidationProblem, ZarrV2ArrayMetadata, - ZarrV2ArrayMetadataPartial, ZarrV2ArrayMetadataStoreKey, + ZarrV2ArrayMetadataUpdate, ZarrV2AttributesStoreKey, ZarrV2ConsolidatedMetadata, ZarrV2ConsolidatedMetadataStoreKey, ZarrV2GroupMetadata, - ZarrV2GroupMetadataPartial, ZarrV2GroupMetadataStoreKey, + ZarrV2GroupMetadataUpdate, ZarrV3ArrayMetadata, - ZarrV3ArrayMetadataPartial, ZarrV3ArrayMetadataStoreKey, + ZarrV3ArrayMetadataUpdate, ZarrV3ConsolidatedMetadata, + ZarrV3ConsolidatedMetadataInput, ZarrV3GroupMetadata, - ZarrV3GroupMetadataPartial, ZarrV3GroupMetadataStoreKey, - ZarrV3MetadataField, - ZarrV3NamedConfig, + ZarrV3GroupMetadataUpdate, + ZarrV3NodeMetadataInput, ) from zarr_metadata.v2.array import ( ZARR_V2_ARRAY_DIMENSION_SEPARATOR, @@ -362,8 +362,8 @@ "ZarrV2ArrayMetadata", "ZarrV2ArrayMetadataJSON", "ZarrV2ArrayMetadataJSONPartial", - "ZarrV2ArrayMetadataPartial", "ZarrV2ArrayMetadataStoreKey", + "ZarrV2ArrayMetadataUpdate", "ZarrV2ArrayOrder", "ZarrV2AttributesStoreKey", "ZarrV2CodecMetadata", @@ -374,28 +374,28 @@ "ZarrV2GroupMetadata", "ZarrV2GroupMetadataJSON", "ZarrV2GroupMetadataJSONPartial", - "ZarrV2GroupMetadataPartial", "ZarrV2GroupMetadataStoreKey", + "ZarrV2GroupMetadataUpdate", "ZarrV2ZArrayJSON", "ZarrV2ZAttrsJSON", "ZarrV2ZGroupJSON", "ZarrV3ArrayMetadata", "ZarrV3ArrayMetadataJSON", "ZarrV3ArrayMetadataJSONPartial", - "ZarrV3ArrayMetadataPartial", "ZarrV3ArrayMetadataStoreKey", + "ZarrV3ArrayMetadataUpdate", "ZarrV3ConsolidatedMetadata", + "ZarrV3ConsolidatedMetadataInput", "ZarrV3ConsolidatedMetadataJSON", "ZarrV3ExtensionField", "ZarrV3GroupMetadata", "ZarrV3GroupMetadataJSON", "ZarrV3GroupMetadataJSONPartial", - "ZarrV3GroupMetadataPartial", "ZarrV3GroupMetadataStoreKey", - "ZarrV3MetadataField", + "ZarrV3GroupMetadataUpdate", "ZarrV3MetadataFieldJSON", - "ZarrV3NamedConfig", "ZarrV3NamedConfigJSON", + "ZarrV3NodeMetadataInput", "ZstdCodecMetadata", "ZstdCodecName", "__version__", diff --git a/packages/zarr-metadata/src/zarr_metadata/_json.py b/packages/zarr-metadata/src/zarr_metadata/_json.py index a5c0a9b6af..dee2807796 100644 --- a/packages/zarr-metadata/src/zarr_metadata/_json.py +++ b/packages/zarr-metadata/src/zarr_metadata/_json.py @@ -9,12 +9,16 @@ from __future__ import annotations +import dataclasses +import json import math -from collections.abc import Mapping, Sequence +from collections.abc import Callable, Iterator, Mapping, Sequence from dataclasses import dataclass -from typing import Literal, TypeGuard, cast, get_args +from types import MappingProxyType +from typing import Final, Literal, TypeGuard, cast, get_args, overload from zarr_metadata._common import JSONValue +from zarr_metadata._sentinel import UNSET ProblemKind = Literal["missing_key", "invalid_type", "invalid_value", "invalid_json", "unknown_key"] """Machine-readable classification of a `ValidationProblem`. @@ -35,20 +39,102 @@ """ +JSON_DEPTH: Final = 256 +"""How many levels of nesting a reader walks. + +A value nested deeper is a problem at the level past the last, so no +document, however deep, takes a reader past what the interpreter allows: +every value the package reads is refined first, and refining stops here. +Every reader, writer and comparison takes one frame for each level, and +`copy.deepcopy`, and `pickle` before Python 3.12, two: so a document at +the cap takes about half of the interpreter's default limit, a thousand +frames, and the rest is the caller's. +""" + +_PAST_THE_LEVELS: Final = f"nested deeper than the {JSON_DEPTH} levels a reader walks" +"""The message of the problem a container past the levels a reader walks is.""" + + +class _Ctx(Mapping[str, JSONValue]): + """What a problem holds as its `ctx`: a copy of its own, arrays as tuples, checked to be JSON when it was made, which nothing edits after. + + Its own type, so a problem built of another's `ctx` -- as `replace` + builds one -- knows it was checked, and skips the check; a mapping of + any other type is checked and copied. + """ + + __slots__ = ("_held",) + + def __init__(self, held: dict[str, JSONValue]) -> None: + self._held = held + + def __getitem__(self, key: str) -> JSONValue: + return self._held[key] + + def __iter__(self) -> Iterator[str]: + return iter(self._held) + + def __len__(self) -> int: + return len(self._held) + + def __repr__(self) -> str: + return repr(self._held) + + def __reduce__(self) -> tuple[type[_Ctx], tuple[dict[str, JSONValue]]]: + return _Ctx, (self._held,) + + +_NO_CTX: Final[Mapping[str, JSONValue]] = _Ctx({}) + + +def _no_ctx() -> Mapping[str, JSONValue]: + """What a problem that says nothing more than its type was expected holds as its `ctx`: nothing.""" + return _NO_CTX + + @dataclass(frozen=True, slots=True) class ValidationProblem: - """A single problem found in a value: where it is, what is wrong, and what kind of wrong. + """A single problem found in a value: where it is, what is wrong, what kind of wrong, and the data the message is made of. `loc` is the path from the root of what was judged to the offending value, e.g. `("codecs", 0, "name")` in a document, and an empty `loc` refers to that root. `kind` classifies the failure mode for programmatic dispatch; `message` is the human-readable description. + + `input` and `ctx` are what the message says, as data, as pydantic's + `ErrorDetails` and zod's issues carry theirs. `input` is the JSON found + at `loc` -- `12`, for a gzip `level` of 12 -- and `UNSET` where nothing + is there, as zod has it for a key that is missing (pydantic gives the + object missing it), or where what is there is not JSON, which the + message shows. It is the object the caller handed in, as pydantic's + is, not a copy: a caller that changes its document afterwards changes + what `input` shows. `ctx` is what was expected, where that is more + than a type: + + - `gt`, `ge`, `lt` and `le`: the bounds the value's type carries, as + pydantic names them -- `{"ge": 0, "le": 9}` for a gzip `level`, + whose type is `Annotated[int, Interval(ge=0, le=9)]` -- or a rule + says. + - `expected`: the values of a closed set, as zod's `values` holds + them, in the order the message lists them -- a `Literal`'s, + `node_type`'s, `zarr_format`'s. + + Neither takes part in equality or the repr: a problem is the same + problem when it is found at the same place and says the same thing. + Every function that returns or raises problems fills `input` from the + value its caller handed it, so a rule says only where a problem is. """ loc: tuple[str | int, ...] message: str kind: ProblemKind + input: JSONValue | UNSET = dataclasses.field( + default=UNSET, kw_only=True, compare=False, repr=False + ) + ctx: Mapping[str, JSONValue] = dataclasses.field( + default_factory=_no_ctx, kw_only=True, compare=False, repr=False + ) def __post_init__(self) -> None: # The runtime half of the annotations: a rule written without a type @@ -69,11 +155,49 @@ def __post_init__(self) -> None: if not isinstance(kind, str) or kind not in get_args(ProblemKind): msg = f"a ValidationProblem's kind is one of {get_args(ProblemKind)!r}, got {kind!r}" raise TypeError(msg) + ctx = cast("object", self.ctx) + if isinstance(ctx, _Ctx): + # A problem's own, checked when it was made: what `replace` + # hands a copy. + return + if ( + not isinstance(ctx, Mapping) + or not all(isinstance(key, str) for key in cast("Mapping[object, object]", ctx)) + or not is_json(dict(cast("Mapping[str, object]", ctx))) + ): + msg = f"a ValidationProblem's ctx is an object of JSON values, got {ctx!r}" + raise TypeError(msg) + # Held as a view of a copy of its own at every level, arrays as + # tuples, so a raised error, a finished report, cannot be edited + # through it, nor through what was handed in. + held = cast( + "dict[str, JSONValue]", + copied(cast("JSONValue", arrays_to_tuples(dict(cast("Mapping[str, object]", ctx))))), + ) + object.__setattr__(self, "ctx", _Ctx(held)) def __str__(self) -> str: location = ".".join(str(part) for part in self.loc) if self.loc else "" return f"{location}: {self.message}" + def __reduce__( + self, + ) -> tuple[Callable[..., ValidationProblem], tuple[object, ...]]: + # Pickled and copied as its constructor called again: the view + # `ctx` is held as does not pickle, and the dict it views does. + return (_problem, (self.loc, self.message, self.kind, self.input, dict(self.ctx))) + + +def _problem( + loc: tuple[str | int, ...], + message: str, + kind: ProblemKind, + found: JSONValue | UNSET, + ctx: Mapping[str, JSONValue], +) -> ValidationProblem: + """A problem built again from what `ValidationProblem.__reduce__` gives.""" + return ValidationProblem(loc, message, kind, input=found, ctx=ctx) + class MetadataValidationError(ValueError): """Raised when a value fails validation, by the entry points that raise rather than report. @@ -114,12 +238,121 @@ def prefixed( loc_head: str | int, problems: Sequence[ValidationProblem] ) -> tuple[ValidationProblem, ...]: """Prepend `loc_head` to the `loc` of every problem (for nested validators).""" - return tuple(ValidationProblem((loc_head, *p.loc), p.message, p.kind) for p in problems) + return tuple(dataclasses.replace(p, loc=(loc_head, *p.loc)) for p in problems) + + +def value_at(value: object, loc: tuple[str | int, ...]) -> object: + """What `value` holds at `loc`, each key naming a member of an object and each index an element of an array; `UNSET` where it holds nothing.""" + for part in loc: + if isinstance(part, str) and is_object(value): + members = value + if part not in members: + return UNSET + value = members[part] + elif isinstance(part, int) and is_array(value) and 0 <= part < len(value): + value = value[part] + else: + return UNSET + return value + + +def with_input( + problems: Sequence[ValidationProblem], value: object, loc: tuple[str | int, ...] = () +) -> tuple[ValidationProblem, ...]: + """`problems`, found in `value`, which sits at `loc`: each given, as its `input`, the JSON `value` holds at its own `loc`. + + What every function that judges a value does with what it found, so a + rule says only where a problem is, and the problem carries what is + there. The function a caller called is the last to do it, so a + problem's `input` is what the value the caller handed in holds, not a + copy some reader inside it made. A problem whose `loc` names nothing + in `value` -- a key that is missing -- or names what is not JSON + keeps what it holds, `UNSET` unless a reader inside found JSON there: + no problem holds what might not pickle, or copy. + """ + if len(problems) == 0: + return () + filled: list[ValidationProblem] = [] + for found in problems: + if found.loc[: len(loc)] == loc: + there = value_at(value, found.loc[len(loc) :]) + if there is not found.input and _holdable(there): + found = dataclasses.replace(found, input=cast("JSONValue", there)) + filled.append(found) + return tuple(filled) + + +def _holdable(value: object) -> bool: + """Whether a problem can hold `value` as its input: JSON, nested no deeper than a reader walks. + + What a validator did not walk -- a member it only reports -- may be + deeper than that; such a value is held as nothing, as + `is_canonical_json` says. + """ + return is_canonical_json(value, finite=False) + + +def not_an_object(value: object) -> tuple[ValidationProblem, ...]: + """What is wrong with a document that is not an object: the one problem, at its root.""" + return with_input((ValidationProblem((), "expected an object", "invalid_type"),), value) + +def validate_json(value: object, loc: tuple[str | int, ...] = ()) -> tuple[ValidationProblem, ...]: + """Return every reason `value`, which sits at `loc`, is not JSON, each where it sits: a float that is not finite, a key that is not a string, a value of no JSON type, a level of nesting past `JSON_DEPTH`, counted from the document's root, which `loc` is below.""" + return with_input(refine_json(value, loc)[1], value, loc) -def validate_json(value: object) -> tuple[ValidationProblem, ...]: - """Return every reason `value` is not JSON-serializable (recursively): `refine_json`'s problems.""" - return refine_json(value)[1] + +def is_object(value: object) -> TypeGuard[Mapping[object, object]]: + """Whether `value` is a JSON object as Python holds one: a mapping, of any keys. + + What `isinstance(value, Mapping)` says, narrowed to a mapping of + `object`, which the bare check leaves unknown to a type checker. + """ + return isinstance(value, Mapping) + + +def refined_object(value: object) -> dict[str, JSONValue]: + """`value`, an object a read found nothing wrong with, refined as user data: the document a model holds. + + `TypeError` when it is not an object, which a read that found no + problem rules out. + """ + refined, _ = refine_user_data(value) + if not isinstance(refined, dict): + msg = f"expected an object a read found nothing wrong with, got {shown(value)}" + raise TypeError(msg) + return refined + + +def object_at(document: Mapping[str, JSONValue], key: str) -> dict[str, JSONValue]: + """The object `document`, refined, holds at `key`, which a read found there; `TypeError` when something else is, which that read rules out.""" + value = document[key] + if not isinstance(value, dict): + msg = f"expected an object at {key!r}, got {shown(value)}" + raise TypeError(msg) + return value + + +def is_json_object(value: object) -> TypeGuard[Mapping[str, object]]: + """Whether `value` is a JSON object with the keys JSON gives one: a mapping whose keys are all strings.""" + return isinstance(value, Mapping) and all( + isinstance(key, str) for key in cast("Mapping[object, object]", value) + ) + + +def is_list_or_tuple(value: object) -> TypeGuard[list[object] | tuple[object, ...]]: + """Whether `value` is a list or a tuple: the two containers canonical JSON holds an array in.""" + return isinstance(value, (list, tuple)) + + +def is_tuple(value: object) -> TypeGuard[tuple[object, ...]]: + """Whether `value` is a tuple, narrowed to a tuple of `object`.""" + return isinstance(value, tuple) + + +def is_array(value: object) -> TypeGuard[Sequence[object]]: + """Whether `value` is a JSON array as Python holds one: a sequence that is not text or bytes.""" + return isinstance(value, Sequence) and not isinstance(value, (str, bytes, bytearray)) def refine_json( @@ -155,25 +388,71 @@ def refine_user_data( _Refined = tuple[JSONValue | None, tuple[ValidationProblem, ...]] +def nested_past_the_levels(value: object, loc: tuple[str | int, ...]) -> ValidationProblem | None: + """The problem `value` is when it is a container -- an object or an array -- at `loc`, past the levels a reader walks, `JSON_DEPTH` of them; None for a scalar, or within them. + + `_refine` asks it of every value it reaches, and a reader of each + container it walks without refining -- a document's members, the + consolidated metadata it descends into -- so a chain of documents is + bounded as any other nesting is, and every container is judged where + it sits, as `refine_json` of the whole document would judge it. + """ + if len(loc) < JSON_DEPTH or isinstance(value, (str, int, float, bool)) or value is None: + return None + if not isinstance(value, (Mapping, Sequence)) or isinstance(value, (bytes, bytearray)): + return None + return ValidationProblem(loc, _PAST_THE_LEVELS, "invalid_value") + + +def within( + problems: Sequence[ValidationProblem], at: tuple[str | int, ...] +) -> tuple[ValidationProblem, ...]: + """`problems`, found in a value that sits at `at`, located from that value: the reverse of `prefixed`, for a reader that counts the levels it walks from the document handed in, but reports where a problem sits in the one it reads. + + A problem not below `at` is a `TypeError`: a reader that located one + from the wrong root would otherwise report it in the wrong place. + """ + for problem in problems: + if problem.loc[: len(at)] != at: + msg = f"a problem at {problem.loc!r} does not sit below {at!r}" + raise TypeError(msg) + return tuple(dataclasses.replace(p, loc=p.loc[len(at) :]) for p in problems) + + def _refine(value: object, loc: tuple[str | int, ...], *, finite: bool) -> _Refined: """`refine_json`, a non-finite number being JSON unless `finite`.""" if isinstance(value, float): if not finite or math.isfinite(value): - return value, () + return float(value), () return None, ( ValidationProblem(loc, f"non-finite float {value!r} is not JSON", "invalid_value"), ) - if isinstance(value, (str, int, bool)) or value is None: + if isinstance(value, int) and not isinstance(value, bool) and not _writable(value): + message = ( + f"an integer of {value.bit_length()} bits has more digits than JSON text holds " + "here, as sys.get_int_max_str_digits bounds it" + ) + return None, (ValidationProblem(loc, message, "invalid_value"),) + if isinstance(value, bool) or value is None: return value, () - if isinstance(value, Mapping): + if isinstance(value, str): + # A subclass -- a `StrEnum` member -- as the string JSON writes for it. + return str.__str__(value), () + if isinstance(value, int): + return int(value), () + if (past := nested_past_the_levels(value, loc)) is not None: + return None, (past,) + if is_object(value): # Walked here rather than through `_refine_members`, so that each - # level of nesting costs one frame, as deep as the interpreter goes. + # level of nesting costs one frame, `JSON_DEPTH` of them at most. members: dict[str, JSONValue] = {} found_in_members: list[ValidationProblem] = [] - for key, item in cast("Mapping[object, object]", value).items(): + for key, item in value.items(): if not isinstance(key, str): found_in_members.append( - ValidationProblem(loc, f"non-string key {key!r} in JSON object", "invalid_type") + ValidationProblem( + loc, f"non-string key {shown_key(key)} in JSON object", "invalid_type" + ) ) continue member, found = _refine(item, (*loc, key), finite=finite) @@ -181,40 +460,169 @@ def _refine(value: object, loc: tuple[str | int, ...], *, finite: bool) -> _Refi if len(found) == 0: members[key] = member return (members if len(found_in_members) == 0 else None), tuple(found_in_members) - if isinstance(value, Sequence) and not isinstance(value, (bytes, bytearray)): + if is_array(value): entries: list[JSONValue] = [] found_in_entries: list[ValidationProblem] = [] - for index, item in enumerate(cast("Sequence[object]", value)): + for index, item in enumerate(value): entry, found = _refine(item, (*loc, index), finite=finite) found_in_entries.extend(found) if len(found) == 0: entries.append(entry) return (tuple(entries) if len(found_in_entries) == 0 else None), tuple(found_in_entries) return None, ( - ValidationProblem(loc, f"not a JSON-serializable value: {value!r}", "invalid_type"), + ValidationProblem( + loc, f"not a JSON-serializable value: {shown_by_python(value)}", "invalid_type" + ), ) +def _writable(value: int) -> bool: + """Whether the interpreter converts `value` to text: what `json.dumps` and `repr` do, which `sys.set_int_max_str_digits` bounds; an integer past the bound could never be written back.""" + if value.bit_length() <= 64: + return True + try: + str(value) + except ValueError: + return False + return True + + +def json_text(value: JSONValue) -> str: + """`value` as the JSON text `json.dumps` writes for it, an object's keys sorted: what `==` compares of a JSON value the package does not interpret. + + Not Python's `==` on the value, which takes `true` for `1`, `-0.0` for + `0.0` and `NaN` for no value at all, but what a document writes: two + values written alike are one value to every reader. + """ + return json.dumps(value, sort_keys=True, ensure_ascii=False, default=_as_object) + + +def _as_object(value: object) -> dict[object, object]: + """A read-only view of an object, as `json.dumps` is handed one, written as the object; anything else is the `TypeError` `json.dumps` raises.""" + if is_object(value): + return dict(value) + msg = f"{value!r} is not JSON" + raise TypeError(msg) + + +def shown(value: object) -> str: + """`value` as a problem's message shows it: as the JSON a document writes, `null` and `[1, 2]`, or by its repr when it is not JSON; what the interpreter will not write, an integer of too many digits or a value nested too deep, by saying so.""" + if isinstance(value, int) and not isinstance(value, bool) and not _writable(value): + return f"an integer of {value.bit_length()} bits" + refined, problems = _refine(value, (), finite=False) + if any(problem.message == _PAST_THE_LEVELS for problem in problems): + # Nested past the levels a reader walks: said so, not left to the + # repr, which overflows at a depth the interpreter and platform set. + return "a value nested too deep to show" + if len(problems) != 0: + # Not JSON. + return shown_by_python(value) + try: + return json.dumps(refined, ensure_ascii=False) + except ValueError: + # An integer of more digits than the interpreter converts to text: + # the value itself, by its size, or one a container holds, which + # is then what Python will not write either. + if isinstance(value, int): + return f"an integer of {value.bit_length()} bits" + return shown_by_python(value) + + +def shown_key(key: object) -> str: + """A key that is not a string, as a message shows it: as Python shows it, since no JSON object holds such a key to write; an integer of more digits than the interpreter writes, by its size.""" + if isinstance(key, int) and not isinstance(key, bool): + try: + return repr(key) + except ValueError: + return f"an integer of {key.bit_length()} bits" + return shown_by_python(key) + + +def shown_by_python(value: object) -> str: + """`value` as Python shows it, for what is not JSON a reader walks; what the interpreter will not write -- holding an integer of more digits than it writes -- said so.""" + try: + return repr(value) + except ValueError: + return f"a value of type {type(value).__name__} the interpreter will not write" + + +def choices(allowed: Sequence[object]) -> str: + """A closed set of values as a message names it: `"C"` alone, or `one of ["C", "F"]`.""" + values = [shown(value) for value in listed(allowed)] + return values[0] if len(values) == 1 else f"one of [{', '.join(values)}]" + + +def listed(allowed: Sequence[object]) -> tuple[object, ...]: + """A closed set of values, each once, in the order a message lists them: by the JSON each is written as.""" + written = {shown(value): value for value in allowed} + return tuple(written[json_text] for json_text in sorted(written)) + + +def outside_of( + loc: tuple[str | int, ...], value: object, allowed: Sequence[object] +) -> ValidationProblem: + """The problem with `value`, found at `loc`, outside the closed set `allowed`. + + Its message lists the set, its kind says whether the value is of the + wrong JSON type or of the right one with the wrong value, and its + `ctx` holds the set, as `expected`. + """ + message = f"expected {choices(allowed)}, got {shown(value)}" + expected = cast("tuple[JSONValue, ...]", listed(allowed)) + return ValidationProblem(loc, message, refused_kind(value, allowed), ctx={"expected": expected}) + + +def json_type(value: object) -> str: + """The JSON type of `value`, as a message names it: "a string", "a number", "null".""" + if value is None: + return "null" + if isinstance(value, bool): + return "a boolean" + if isinstance(value, (int, float)): + return "a number" + if isinstance(value, str): + return "a string" + if isinstance(value, Mapping): + return "an object" + if isinstance(value, (list, tuple)): + return "an array" + return "a value" + + +def refused_kind(value: object, allowed: Sequence[object]) -> ProblemKind: + """What is wrong with `value`, outside a closed set: its type, when none of the set is of its JSON type, else its value.""" + same_type = json_type(value) in {json_type(entry) for entry in allowed} + return "invalid_value" if same_type else "invalid_type" + + def is_canonical_json(value: object, *, finite: bool = True) -> TypeGuard[JSONValue]: - """Whether `value` already uses the concrete containers in `JSONValue`. + """Whether `value` already uses the concrete containers in `JSONValue`, nested no deeper than a reader walks. A non-finite number counts only when `finite` is false, as a document's guard passes it: where one may be is the document's validator's to say. - One frame per level of nesting, as `refine_json` takes, so a value - `refine_json` reads is one this can walk. + One frame per level of nesting, `JSON_DEPTH` of them at most, as + `refine_json` takes: a container past them is no JSON a reader walks, + so it is none to this either, and the walk stops there. """ + return _is_canonical(value, 0, finite=finite) + + +def _is_canonical(value: object, depth: int, *, finite: bool) -> bool: + """`is_canonical_json` of `value`, which sits `depth` levels down.""" if isinstance(value, float): return not finite or math.isfinite(value) if isinstance(value, (str, int, bool)) or value is None: return True - if isinstance(value, (list, tuple)): - for item in cast("list[object] | tuple[object, ...]", value): - if not is_canonical_json(item, finite=finite): + if depth >= JSON_DEPTH: + return False + if is_list_or_tuple(value): + for item in value: # noqa: SIM110 - a loop, not a generator, is one frame per level + if not _is_canonical(item, depth + 1, finite=finite): return False return True - if isinstance(value, dict): - for key, item in cast("dict[object, object]", value).items(): - if not isinstance(key, str) or not is_canonical_json(item, finite=finite): + if is_object(value): + for key, item in value.items(): + if not isinstance(key, str) or not _is_canonical(item, depth + 1, finite=finite): return False return True return False @@ -229,26 +637,66 @@ def parse_json(value: object) -> JSONValue: """Return a canonical `JSONValue`, or raise `MetadataValidationError`.""" refined, problems = refine_json(value) if len(problems) != 0: - raise MetadataValidationError(problems) + raise MetadataValidationError(with_input(problems, value)) return refined +@overload +def frozen(value: Mapping[str, JSONValue]) -> Mapping[str, JSONValue]: ... +@overload +def frozen(value: JSONValue) -> JSONValue: ... +def frozen(value: JSONValue) -> JSONValue: + """`value` as a read-only view at every level: each object a mapping proxy, each array a tuple, so nothing handed out can be changed in place. + + Shares the scalars with `value`, and copies nothing else than the + containers a view needs. What a model shows of its document. + """ + if isinstance(value, Mapping): + return cast( + "JSONValue", MappingProxyType({key: frozen(item) for key, item in value.items()}) + ) + if isinstance(value, (tuple, list)): + return tuple(frozen(item) for item in value) + return value + + +def copied(value: JSONValue) -> JSONValue: + """`value` in containers of its own, sharing nothing with it: each object a new `dict`, each array a new one of its type. + + One frame for each level of nesting, as `refine_json` reads, so a + value refined is copied however deep it is. + """ + if isinstance(value, Mapping): + members: dict[str, JSONValue] = {} + for key, item in value.items(): + members[key] = copied(item) + return members + if isinstance(value, (tuple, list)): + # A loop, not a comprehension, which is a frame of its own before + # Python 3.12: one frame for each level. + entries = list(map(copied, value)) + return entries if isinstance(value, list) else tuple(entries) + return value + + def arrays_to_tuples(obj: object) -> object: """Recursively materialize mappings and convert array-like values to tuples.""" - if isinstance(obj, Sequence) and not isinstance(obj, (str, bytes, bytearray)): - sequence = cast("Sequence[object]", obj) - converted_sequence = tuple(arrays_to_tuples(item) for item in sequence) + if is_array(obj): + sequence = obj + # Loops, not comprehensions, which are a frame of their own before + # Python 3.12: one frame for each level, as `copied` takes. + converted_sequence = tuple(map(arrays_to_tuples, sequence)) if isinstance(obj, tuple) and all( converted is original for converted, original in zip(converted_sequence, sequence, strict=True) ): return sequence return converted_sequence - if isinstance(obj, Mapping): - mapping = cast("Mapping[object, object]", obj) - converted: dict[object, object] = { - key: arrays_to_tuples(value) for key, value in mapping.items() - } + if is_object(obj): + mapping = obj + converted: dict[object, object] = {} + for key, value in mapping.items(): + converted[key] = arrays_to_tuples(value) if isinstance(obj, dict) and all(converted[key] is value for key, value in mapping.items()): return mapping return converted @@ -260,11 +708,25 @@ def arrays_to_tuples(obj: object) -> object: "ProblemKind", "ValidationProblem", "arrays_to_tuples", + "choices", + "copied", + "frozen", "is_canonical_json", "is_json", + "json_type", + "listed", + "nested_past_the_levels", + "not_an_object", + "outside_of", "parse_json", "prefixed", "refine_json", "refine_user_data", + "refused_kind", + "shown", + "shown_key", "validate_json", + "value_at", + "with_input", + "within", ] diff --git a/packages/zarr-metadata/src/zarr_metadata/_pydantic_schema.py b/packages/zarr-metadata/src/zarr_metadata/_pydantic_schema.py index 8b0e354bc4..64e4d01bd5 100644 --- a/packages/zarr-metadata/src/zarr_metadata/_pydantic_schema.py +++ b/packages/zarr-metadata/src/zarr_metadata/_pydantic_schema.py @@ -2,27 +2,36 @@ from __future__ import annotations -from collections.abc import Mapping # noqa: TC003 # resolved by Pydantic at runtime +import re +from collections.abc import Mapping # resolved by Pydantic at runtime from typing import Annotated, Literal, NotRequired from pydantic import Field from typing_extensions import TypedDict from zarr_metadata._common import JSONValue -from zarr_metadata.v2.array import ( # noqa: TC001 # resolved by Pydantic at runtime +from zarr_metadata.v2.array import ( # resolved by Pydantic at runtime ZarrV2DataTypeMetadata, ) from zarr_metadata.v2.codec import ( # resolved by Pydantic at runtime ZarrV2CodecMetadata, ) +from zarr_metadata.v3._common import EXTENSION_NAME_SCHEMA_PATTERN NonNegativeInt = Annotated[int, Field(ge=0)] +ExtensionName = Annotated[str, Field(pattern=re.compile(EXTENSION_NAME_SCHEMA_PATTERN))] +"""A name as the spec names an extension, the pattern `well_named` accepts, so the schema pydantic generates refuses what the reader refuses. + +Compiled, so pydantic reads it with Python's `re`, which has the look-ahead +the pattern ends in where its own engine has none, and writes it into the +schema as it is. +""" class ZarrV3NamedConfigJSON(TypedDict, closed=True): """Closed v3 named configuration read on its own, outside a document, where `must_understand` may be `false`.""" - name: str + name: ExtensionName configuration: NotRequired[Mapping[str, JSONValue]] must_understand: NotRequired[bool] @@ -30,13 +39,13 @@ class ZarrV3NamedConfigJSON(TypedDict, closed=True): class ZarrV3MandatoryNamedConfigJSON(TypedDict, closed=True): """Closed named configuration at an extension point of a document, where understanding is mandatory.""" - name: str + name: ExtensionName configuration: NotRequired[Mapping[str, JSONValue]] must_understand: NotRequired[Literal[True]] -ZarrV3MetadataFieldJSON = str | ZarrV3NamedConfigJSON -ZarrV3MandatoryMetadataFieldJSON = str | ZarrV3MandatoryNamedConfigJSON +ZarrV3MetadataFieldJSON = ExtensionName | ZarrV3NamedConfigJSON +ZarrV3MandatoryMetadataFieldJSON = ExtensionName | ZarrV3MandatoryNamedConfigJSON ZarrV3CodecPipelineJSON = Annotated[ tuple[ZarrV3MandatoryMetadataFieldJSON, ...], Field(min_length=1) ] @@ -73,7 +82,7 @@ class ZarrV3GroupMetadataJSON(TypedDict, extra_items=JSONValue): zarr_format: Literal[3] node_type: Literal["group"] attributes: NotRequired[Mapping[str, JSONValue]] - consolidated_metadata: NotRequired[ZarrV3ConsolidatedMetadataJSON | None] + consolidated_metadata: NotRequired[ZarrV3ConsolidatedMetadataJSON] class ZarrV2ArrayMetadataJSON(TypedDict, extra_items=JSONValue): diff --git a/packages/zarr-metadata/src/zarr_metadata/model/_sentinel.py b/packages/zarr-metadata/src/zarr_metadata/_sentinel.py similarity index 90% rename from packages/zarr-metadata/src/zarr_metadata/model/_sentinel.py rename to packages/zarr-metadata/src/zarr_metadata/_sentinel.py index ad71e216fa..7b159688b2 100644 --- a/packages/zarr-metadata/src/zarr_metadata/model/_sentinel.py +++ b/packages/zarr-metadata/src/zarr_metadata/_sentinel.py @@ -4,7 +4,9 @@ JSON `null` in the document (a v2 `compressor`/`filters` value, an unnamed dimension inside `dimension_names`), and `UNSET` always means the document key is absent. The two are never interchangeable, so a model value can never -leak into a document as a spelling the writer did not intend. +leak into a document as a spelling the writer did not intend. A problem's +`input` keeps the same invariant: it is `UNSET` where nothing was found at +the problem's `loc`, and `None` where a `null` was. Check with identity: `if model.dimension_names is UNSET: ...`. diff --git a/packages/zarr-metadata/src/zarr_metadata/_typed_json.py b/packages/zarr-metadata/src/zarr_metadata/_typed_json.py index b8dbc1cd80..3d7bdd0f2c 100644 --- a/packages/zarr-metadata/src/zarr_metadata/_typed_json.py +++ b/packages/zarr-metadata/src/zarr_metadata/_typed_json.py @@ -8,7 +8,10 @@ `tuple[T1, T2]`, a union of those, an object described by a TypedDict, `Mapping[str, V]`, a `NewType` as the type it names, and a type alias as the type it stands for -- which is what keeps it small. An annotation -outside these implies no parser. +outside these implies no parser. A number's type may carry bounds, as +annotated-types spells them and pydantic reads them -- +`Annotated[int, Interval(ge=0, le=9)]` -- and a value out of them is a +problem; `constraints_of` says which the checker reads. A TypedDict is read as the typing spec defines it, which is not always what its runtime attributes say: `typeddict_keys` reads which keys it @@ -25,15 +28,22 @@ depth; the parser it returns is used as it is. Parsers are compiled once per annotation and are pure functions of the value, so the branch of a union that did not match leaves nothing behind. + +`json_schema` writes the same reading as a JSON Schema: `Schemas` writes +each shape as the checker reads it, asking a caller's `SchemaLeaf` first, +as a parser asks a `Leaf`. """ from __future__ import annotations import functools +import math +import operator import sys import types import typing -from collections.abc import Callable, Mapping, Sequence +import urllib.parse +from collections.abc import Callable, Iterator, Mapping, Sequence from collections.abc import Set as AbstractSet from dataclasses import dataclass from typing import ( @@ -46,6 +56,7 @@ NewType, NoReturn, TypeAlias, + TypeGuard, TypeVar, cast, get_args, @@ -53,11 +64,23 @@ get_type_hints, ) +import annotated_types import typing_extensions from typing_extensions import NoExtraItems, TypeIs, is_typeddict from zarr_metadata._common import JSONValue -from zarr_metadata._json import ValidationProblem, is_json, refine_json +from zarr_metadata._json import ( + ValidationProblem, + choices, + copied, + is_json, + is_list_or_tuple, + is_object, + outside_of, + refine_json, + shown, + with_input, +) if TYPE_CHECKING: from zarr_metadata._json import ProblemKind @@ -89,9 +112,12 @@ # A `type` statement makes a `typing.TypeAliasType`, which is not the # `typing_extensions` one on every version that has both. -_ALIASES: Final[tuple[type, ...]] = ( +_ALIASES: Final[tuple[type[typing_extensions.TypeAliasType], ...]] = ( typing_extensions.TypeAliasType, - getattr(typing, "TypeAliasType", typing_extensions.TypeAliasType), + cast( + "type[typing_extensions.TypeAliasType]", + getattr(typing, "TypeAliasType", typing_extensions.TypeAliasType), + ), ) @@ -131,7 +157,9 @@ def strip_annotation(annotation: object) -> tuple[object, tuple[object, ...]]: origin = get_origin(annotation) if origin is Annotated: inner, *extras = get_args(annotation) - metadata.extend(extras) + # An inner layer's metadata first, as `Annotated` flattens + # nested layers. + metadata[:0] = extras annotation = inner elif origin in _QUALIFIERS: (annotation,) = get_args(annotation) @@ -139,6 +167,14 @@ def strip_annotation(annotation: object) -> tuple[object, tuple[object, ...]]: return annotation, tuple(metadata) +def unqualified(annotation: object) -> object: + """A TypedDict key's annotation with its qualifiers peeled and its `Annotated` metadata kept, which says more of the value's type.""" + inner, metadata = strip_annotation(annotation) + if len(metadata) == 0: + return inner + return Annotated[(inner, *metadata)] + + Qualifier: TypeAlias = Literal["Required", "NotRequired", "ReadOnly"] """A qualifier on a TypedDict key, by name: `typing` and `typing_extensions` may each spell one.""" @@ -166,11 +202,9 @@ def is_union(annotation: object) -> bool: return get_origin(annotation) in (typing.Union, types.UnionType) -def is_alias(annotation: object) -> bool: +def is_alias(annotation: object) -> TypeGuard[typing_extensions.TypeAliasType]: """Whether `annotation` is a type alias that takes no type parameters: `type Level = int`.""" - if not isinstance(annotation, _ALIASES): - return False - return len(cast("typing_extensions.TypeAliasType", annotation).__type_params__) == 0 + return isinstance(annotation, _ALIASES) and len(annotation.__type_params__) == 0 @functools.cache @@ -202,10 +236,11 @@ class TypedDictKeys: """What a TypedDict says of an object's keys, read as the typing spec defines it. `members` holds every key it declares, its bases' included, with the - type of the key's value -- qualifiers and `Annotated` metadata peeled - -- and whether the key is required. `extra_items` is what any other - key may hold: `Never` when the TypedDict is closed, `object` when it - is open, and the `extra_items` type otherwise. `declared` is whether + type of the key's value -- its qualifiers peeled, and any `Annotated` + metadata kept, since a constraint there is part of the type -- and + whether the key is required. `extra_items` is what any other key may + hold: `Never` when the TypedDict is closed, `object` when it is open, + and the `extra_items` type otherwise. `declared` is whether that was said, by the TypedDict or a base, rather than defaulted: a TypedDict that says nothing is open. """ @@ -222,12 +257,13 @@ def required(self) -> frozenset[str]: @property def closed(self) -> bool: """Whether a key it does not declare is not a key of the type: `closed=True`, or `extra_items=Never`.""" - return self.extra_items is Never or self.extra_items is NoReturn + extra_items = strip_annotation(self.extra_items)[0] + return extra_items is Never or extra_items is NoReturn @property def open(self) -> bool: """Whether a key it does not declare may hold anything: the default, or `closed=False`.""" - return self.extra_items is object + return strip_annotation(self.extra_items)[0] is object @functools.cache @@ -269,7 +305,7 @@ def typeddict_keys(typeddict: type) -> TypedDictKeys: required = False else: required = key in required_at_runtime - members[key] = (strip_annotation(hint)[0], required) + members[key] = (unqualified(hint), required) extra_items, declared = _openness(typeddict) return TypedDictKeys(types.MappingProxyType(members), extra_items, declared) @@ -332,7 +368,7 @@ def _openness(typeddict: type) -> tuple[object, bool]: if closed is not None: return (Never if closed else object), True inherited = [_openness(base) for base in _typeddict_bases(typeddict)] - restricted = {extra for extra, _ in inherited if extra is not object} + restricted = {extra for extra, _ in inherited if strip_annotation(extra)[0] is not object} if len(restricted) > 1: msg = ( f"{typeddict.__name__}: its bases disagree on what a key they do not declare may " @@ -345,7 +381,7 @@ def _openness(typeddict: type) -> tuple[object, bool]: def _extra_items_type(typeddict: type, extra_items: object) -> object: - """The `extra_items` type, evaluated where `typeddict` was defined, `ReadOnly` peeled.""" + """The `extra_items` type, evaluated where `typeddict` was defined, `ReadOnly` peeled and any `Annotated` metadata kept.""" if isinstance(extra_items, (str, ForwardRef)): holder = type( "_ExtraItems", @@ -357,7 +393,7 @@ def _extra_items_type(typeddict: type, extra_items: object) -> object: except NameError as error: msg = f"{typeddict.__name__}: {error}; its extra_items must resolve in its module" raise TypeError(msg) from error - return strip_annotation(extra_items)[0] + return unqualified(extra_items) def _typeddict_bases(typeddict: type) -> tuple[type, ...]: @@ -375,43 +411,66 @@ def _typeddict_bases(typeddict: type) -> tuple[type, ...]: def describe(annotation: object, seen: frozenset[object] = frozenset()) -> str: """The annotation as a message would name it: "an integer", "an object".""" + return _named(annotation, seen)[0] + + +def _named(annotation: object, seen: frozenset[object]) -> tuple[str, str]: + """The annotation as a message names one of it and many of it: "an integer", "integers".""" inner = strip_annotation(annotation)[0] if inner is int: - return "an integer" + return "an integer", "integers" if inner is float: - return "a number" + return "a number", "numbers" if inner is bool: - return "a boolean" + return "a boolean", "booleans" if inner is str: - return "a string" + return "a string", "strings" if inner is None or inner is types.NoneType: - return "null" + return "null", "nulls" if inner is JSONValue: - return "a JSON value" + return "a JSON value", "JSON values" origin = get_origin(inner) if origin is Literal: - return f"one of {tuple(sorted(get_args(inner), key=repr))!r}" + values = get_args(inner) + listed = ", ".join(sorted(dict.fromkeys(shown(value) for value in values))) + return choices(values), f"values in [{listed}]" if is_union(inner): - return " or ".join(describe(branch, seen) for branch in get_args(inner)) + # Each shape once -- two TypedDicts are both "an object" -- and a + # broader one takes in a narrower: a number an integer, a string + # the strings a `Literal` names. + branches = get_args(inner) + named = dict.fromkeys(_named(branch, seen) for branch in branches) + if ("a number", "numbers") in named: + named.pop(("an integer", "integers"), None) + if ("a string", "strings") in named: + for branch in filter(_strings_only, branches): + named.pop(_named(branch, seen), None) + return " or ".join(one for one, _ in named), " or ".join(many for _, many in named) if origin is tuple: arguments = get_args(inner) if len(arguments) == 2 and arguments[1] is Ellipsis: - return f"an array of {describe(arguments[0], seen)} elements" + each = _named(arguments[0], seen)[1] + return f"an array of {each}", f"arrays of {each}" if len(arguments) == 2: - return f"a [{describe(arguments[0], seen)}, {describe(arguments[1], seen)}] pair" - return f"an array of {len(arguments)} elements" - if origin in (Mapping, dict): - return "an object" + pair = f"[{describe(arguments[0], seen)}, {describe(arguments[1], seen)}] pair" + return f"a {pair}", f"{pair}s" + return f"an array of {len(arguments)} elements", f"arrays of {len(arguments)} elements" + if origin in (Mapping, dict) or (isinstance(inner, type) and is_typeddict(inner)): + return "an object", "objects" if isinstance(inner, NewType): - return describe(inner.__supertype__, seen) - if isinstance(inner, type) and is_typeddict(inner): - return "an object" + return _named(inner.__supertype__, seen) if is_alias(inner): - alias = cast("typing_extensions.TypeAliasType", inner) + alias = inner if alias in seen or _holds(alias_value(alias), alias, frozenset()): - return f"a {alias.__name__}" - return describe(alias_value(alias), seen | {alias}) - return "a value" + return f"a {alias.__name__}", f"{alias.__name__} values" + return _named(alias_value(alias), seen | {alias}) + return "a value", "values" + + +def _strings_only(annotation: object) -> bool: + """Whether `annotation` is a `Literal` of strings, which "a string" takes in.""" + inner = strip_annotation(annotation)[0] + return get_origin(inner) is Literal and all(isinstance(value, str) for value in get_args(inner)) def _holds(annotation: object, alias: object, seen: frozenset[object]) -> bool: @@ -422,7 +481,7 @@ def _holds(annotation: object, alias: object, seen: frozenset[object]) -> bool: if is_alias(inner): if inner in seen: return False - value = alias_value(cast("typing_extensions.TypeAliasType", inner)) + value = alias_value(inner) return _holds(value, alias, seen | {inner}) return any(_holds(argument, alias, seen) for argument in get_args(inner)) @@ -440,7 +499,7 @@ def shape_of(annotation: object) -> str | None: if isinstance(inner, NewType): inner = strip_annotation(inner.__supertype__)[0] else: - inner = strip_annotation(alias_value(cast("typing_extensions.TypeAliasType", inner)))[0] + inner = strip_annotation(alias_value(inner))[0] if inner is int: return "int" if inner is float: @@ -491,7 +550,7 @@ def _scalar(description: str, admits: Callable[[object], bool]) -> Parser: def parse(value: object, loc: Loc) -> Parsed: if admits(value): return value, () - return value, problem(loc, f"expected {description}, got {value!r}") + return value, problem(loc, f"expected {description}, got {shown(value)}") return parse @@ -510,14 +569,14 @@ def one_of(allowed: tuple[object, ...]) -> Parser: """A member whose type is a closed set of values. Equal and of the same type: JSON `true` is not the integer 1, though - Python says `True == 1`. + Python says `True == 1`. A value of a JSON type none of them has -- + a number, where each is a string -- is of the wrong type; one of the + right type, the wrong value. """ def parse(value: object, loc: Loc) -> Parsed: if not any(value == entry and type(value) is type(entry) for entry in allowed): - return value, problem( - loc, f"expected one of {allowed!r}, got {value!r}", "invalid_value" - ) + return value, (outside_of(loc, value, allowed),) return value, () return parse @@ -527,9 +586,9 @@ def sequence_of(element: Parser) -> Parser: """A member whose type is an array of one element type, parsed element by element.""" def parse(value: object, loc: Loc) -> Parsed: - if not isinstance(value, (list, tuple)): - return value, problem(loc, f"expected a sequence, got {value!r}") - entries = cast("list[object] | tuple[object, ...]", value) + if not is_list_or_tuple(value): + return value, problem(loc, f"expected an array, got {shown(value)}") + entries = value parsed: list[object] = [] found: list[ValidationProblem] = [] for index, entry in enumerate(entries): @@ -545,11 +604,11 @@ def fixed_tuple(elements: Sequence[Parser], description: str) -> Parser: """A member whose type is an array of a fixed length, parsed position by position.""" def parse(value: object, loc: Loc) -> Parsed: - if not isinstance(value, (list, tuple)): - return value, problem(loc, f"expected {description}, got {value!r}") - entries = tuple(cast("list[object] | tuple[object, ...]", value)) + if not is_list_or_tuple(value): + return value, problem(loc, f"expected {description}, got {shown(value)}") + entries = tuple(value) if len(entries) != len(elements): - return entries, problem(loc, f"expected {description}, got {entries!r}") + return entries, problem(loc, f"expected {description}, got {shown(entries)}") parsed: list[object] = [] found: list[ValidationProblem] = [] for position, (element, entry) in enumerate(zip(elements, entries, strict=True)): @@ -586,7 +645,7 @@ def any_of(branches: Sequence[Branch], description: str, tag: Tag | None = None) def parse(value: object, loc: Loc) -> Parsed: present = _keys_of(value) if tag is not None and present is not None: - return _by_tag(branches, tag, cast("Mapping[str, object]", value), loc) + return _by_tag(branches, tag, value, loc) clean: list[tuple[int, int, Parsed]] = [] failed: list[tuple[int, int, Parsed]] = [] for index, (shape, branch, keys) in enumerate(branches): @@ -603,30 +662,27 @@ def parse(value: object, loc: Loc) -> Parsed: return min(clean, key=lambda found: found[:2])[2] if len(failed) != 0: return min(failed, key=lambda found: found[:2])[2] - return value, problem(loc, f"expected {description}, got {value!r}") + return value, problem(loc, f"expected {description}, got {shown(value)}") return parse -def _by_tag(branches: Sequence[Branch], tag: Tag, value: Mapping[str, object], loc: Loc) -> Parsed: +def _by_tag(branches: Sequence[Branch], tag: Tag, value: object, loc: Loc) -> Parsed: """`value` parsed by the branch its tag picks; a tag missing, or one no branch has, reported at it.""" key, picks = tag - if key not in value: + if not is_object(value) or key not in value: return value, problem((*loc, key), f"missing required key {key!r}", "missing_key") said = value[key] index = picks.get((type(said), said)) if _hashable(said) else None if index is None: - allowed = tuple(sorted((entry for _, entry in picks), key=repr)) - return value, problem( - (*loc, key), f"expected one of {allowed!r}, got {said!r}", "invalid_value" - ) + return value, (outside_of((*loc, key), said, tuple(entry for _, entry in picks)),) return branches[index][1](value, loc) def _keys_of(value: object) -> AbstractSet[str] | None: """The keys of `value` if it is a JSON object, else None.""" - if isinstance(value, Mapping): - return cast("Mapping[str, object]", value).keys() + if is_object(value): + return {key for key in value if isinstance(key, str)} return None @@ -645,12 +701,15 @@ def object_of(members: Mapping[str, tuple[Parser, bool]], extra: Parser | None) """ def parse(value: object, loc: Loc) -> Parsed: - if not isinstance(value, Mapping): - return value, problem(loc, f"expected an object, got {value!r}") - entries = cast("Mapping[str, object]", value) + if not is_object(value): + return value, problem(loc, f"expected an object, got {shown(value)}") + entries = value parsed: dict[str, object] = {} found: list[ValidationProblem] = [] for key, entry in entries.items(): + if not isinstance(key, str): + found.extend(problem(loc, f"non-string key {shown(key)}")) + continue if key in members: continue if extra is None: @@ -664,7 +723,9 @@ def parse(value: object, loc: Loc) -> Parsed: found.extend(problems) elif required: found.extend(problem((*loc, key), f"missing required key {key!r}", "missing_key")) - return {key: parsed[key] for key in entries if key in parsed}, tuple(found) + return { + key: parsed[key] for key in entries if isinstance(key, str) and key in parsed + }, tuple(found) return parse @@ -677,12 +738,15 @@ def mapping_of(value: Parser) -> Parser: """ def parse(candidate: object, loc: Loc) -> Parsed: - if not isinstance(candidate, Mapping): - return candidate, problem(loc, f"expected an object, got {candidate!r}") - entries = cast("Mapping[str, object]", candidate) + if not is_object(candidate): + return candidate, problem(loc, f"expected an object, got {shown(candidate)}") + entries = candidate parsed: dict[str, object] = {} found: list[ValidationProblem] = [] for key, entry in entries.items(): + if not isinstance(key, str): + found.extend(problem(loc, f"non-string key {shown(key)}")) + continue item, problems = value(entry, (*loc, key)) parsed[key] = item found.extend(problems) @@ -691,6 +755,137 @@ def parse(candidate: object, loc: Loc) -> Parsed: return parse +# --- constraints --------------------------------------------------------- + +Constraints: TypeAlias = Mapping[str, int | float] +"""The bounds a type carries, by the names pydantic gives them: `{"ge": 0, "le": 9}`.""" + +_BOUNDS: Final[tuple[tuple[type, str], ...]] = ( + (annotated_types.Gt, "gt"), + (annotated_types.Ge, "ge"), + (annotated_types.Lt, "lt"), + (annotated_types.Le, "le"), +) +"""The annotated-types constraints the checker reads, each with its name.""" + +_FROM_BELOW: Final = frozenset({"gt", "ge"}) +"""The bounds a value must be above.""" + +_FROM_ABOVE: Final = frozenset({"lt", "le"}) +"""The bounds a value must be below.""" + +_EXCLUSIVE: Final = frozenset({"gt", "lt"}) +"""The bounds a value may not equal.""" + +_HOLDS: Final[Mapping[str, Callable[[float, float], bool]]] = { + "gt": operator.gt, + "ge": operator.ge, + "lt": operator.lt, + "le": operator.le, +} +"""Whether a number keeps within a bound of each name.""" + + +def constraints_of(metadata: Sequence[object]) -> dict[str, int | float]: + """The bounds among an `Annotated` type's metadata, by name: at most one from below, and one from above. + + annotated-types' `Gt`, `Ge`, `Lt` and `Le`, and an `Interval`, which + unpacks to them, as pydantic reads them. A note -- a string, a `Doc` + -- says nothing of what a value is, and is passed over. Anything else + is a `TypeError`: a constraint the checker does not read -- a + `MinLen`, a `Predicate`, pydantic's `Field` -- since a type that says + what its values are, and a checker that does not hold them to it, + would disagree; a second bound from one side, which pydantic reads as + the last one said; and a bound that is not a finite number. A bound + is held as the number it equals, an integer when it is one: + `Ge(True)`, `Ge(1.0)` and `Ge(1)` are one bound, as Python's + `Annotated` cache, which takes equal metadata for the same, may hand + back any of them for another. + """ + said: dict[str, int | float] = {} + for item in _unpacked(metadata): + if isinstance(item, (str, typing_extensions.Doc)): + continue + name = next((name for kind, name in _BOUNDS if isinstance(item, kind)), None) + if name is None: + msg = f"{item!r} is not a constraint the checker reads, which are Gt, Ge, Lt and Le" + raise TypeError(msg) + side = _FROM_BELOW if name in _FROM_BELOW else _FROM_ABOVE + if not side.isdisjoint(said): + where = "below" if side is _FROM_BELOW else "above" + msg = f"{item!r} is a second bound from {where}; a type takes one from each side" + raise TypeError(msg) + said[name] = _bound(item, name) + return said + + +def _unpacked(metadata: Sequence[object]) -> Iterator[object]: + """`metadata`, each group of constraints in it -- an `Interval` -- unpacked.""" + for item in metadata: + if isinstance(item, annotated_types.GroupedMetadata): + yield from _unpacked(tuple(cast("Sequence[object]", item))) + else: + yield item + + +def _bound(item: object, name: str) -> int | float: + """The number `item`, a bound of `name`, holds a value to, an integer when it is one; `TypeError` for what is not a finite number.""" + bound = cast("object", getattr(item, name)) + if isinstance(bound, int): + return int(bound) # a `bool` as the integer it equals, as the `Annotated` cache has it + if isinstance(bound, float) and math.isfinite(bound): + return int(bound) if bound.is_integer() else bound + msg = f"{item!r}: a bound is a finite number" + raise TypeError(msg) + + +def _constrainable(inner: object, said: Constraints) -> None: + """Refuse bounds on what is not a number: a string, an array, a value of any type.""" + if len(said) != 0 and shape_of(inner) not in ("int", "number"): + msg = f"{', '.join(said)}: a bound is on a number, and {describe(inner)} is not one" + raise TypeError(msg) + + +def constrained(parse: Parser, description: str, said: Constraints) -> Parser: + """What `parse` reads, held to the bounds its type carries: one problem, `invalid_value`, when it breaks any. + + The message says what the type admits -- "expected an integer in [0, + 9], got 12" -- and the problem's `ctx` holds the bounds. Only a + value its type reads is held to them, as rules are asked only of a + value of their type. + """ + expected = f"expected {description} {_admitted(said)}" + ctx: dict[str, JSONValue] = dict(said) + holds = tuple((_HOLDS[name], bound) for name, bound in said.items()) + + def parse_constrained(value: object, loc: Loc) -> Parsed: + typed, found = parse(value, loc) + if len(found) != 0: + return typed, found + number = cast("float", typed) + for keeps, bound in holds: + if not keeps(number, bound): + message = f"{expected}, got {shown(value)}" + return typed, (ValidationProblem(loc, message, "invalid_value", ctx=ctx),) + return typed, () + + return parse_constrained + + +def _admitted(said: Constraints) -> str: + """What the bounds admit, as a message says it after the type: "in [0, 9]", "in (0, 1)", ">= 1".""" + low = next(((name, bound) for name, bound in said.items() if name in _FROM_BELOW), None) + high = next(((name, bound) for name, bound in said.items() if name in _FROM_ABOVE), None) + if low is not None and high is not None: + opening = "(" if low[0] in _EXCLUSIVE else "[" + closing = ")" if high[0] in _EXCLUSIVE else "]" + return f"in {opening}{shown(low[1])}, {shown(high[1])}{closing}" + if low is not None: + return f"{'>' if low[0] in _EXCLUSIVE else '>='} {shown(low[1])}" + name, bound = cast("tuple[str, int | float]", high) + return f"{'<' if name in _EXCLUSIVE else '<='} {shown(bound)}" + + # --- the compiler -------------------------------------------------------- _Building: TypeAlias = dict[object, Parser] @@ -797,6 +992,9 @@ def compile_() -> Parser | None: if keys.closed: return object_of(members, None) if keys.open: + # Any JSON value, which a note does not change and a bound + # cannot: vetted all the same. + _constrainable(object, constraints_of(strip_annotation(keys.extra_items)[1])) return object_of(members, _JSON) extra = _compile(keys.extra_items, leaf, building) return None if extra is None else object_of(members, extra) @@ -824,7 +1022,19 @@ def _literal(inner: object) -> Parser | None: def _compile(annotation: object, leaf: Leaf, building: _Building) -> Parser | None: - inner = strip_annotation(annotation)[0] + inner, metadata = strip_annotation(annotation) + parse = _compile_type(inner, leaf, building) + if parse is None or len(metadata) == 0: + return parse + said = constraints_of(metadata) + if len(said) == 0: + return parse + _constrainable(inner, said) + return constrained(parse, describe(inner), said) + + +def _compile_type(inner: object, leaf: Leaf, building: _Building) -> Parser | None: + """The parser of a type, `Annotated` peeled from it.""" found = leaf(inner) if found is not None: return found @@ -856,7 +1066,7 @@ def _compile(annotation: object, leaf: Leaf, building: _Building) -> Parser | No # the code's, for a value it has vouched for. return _compile(inner.__supertype__, leaf, building) if is_alias(inner): - return _alias(cast("typing_extensions.TypeAliasType", inner), leaf, building) + return _alias(inner, leaf, building) return None @@ -929,40 +1139,291 @@ def check( located under `loc`. What comes back holds what `shape` admits and nothing else: a key a closed TypedDict does not declare is reported, as `unknown_key`, and left out, and the value still comes back. - Anything else wrong and it does not. `TypeError` for a `shape` that is - not a TypedDict, or holds something no parser reads. + Anything else wrong and it does not. The levels a reader walks are + counted from the root of the document `loc` places `value` in, so a + `loc` of 255 levels leaves one. `TypeError` for a `shape` that is not + a TypedDict, or holds something no parser reads. """ if not is_typeddict(shape): msg = f"{shape!r} is not a TypedDict" raise TypeError(msg) refined, problems = refine_json(value, loc) if len(problems) != 0: - return None, problems + return None, with_input(problems, value, loc) typed, found = _checker(shape)(refined, loc) readable = all(problem.kind == "unknown_key" for problem in found) - return (cast("T", typed) if readable else None), found + return (cast("T", typed) if readable else None), with_input(found, value, loc) + + +# --- JSON Schema --------------------------------------------------------- + +JSONSchema: TypeAlias = dict[str, JSONValue] +"""A JSON Schema, as the JSON object it is: arrays as lists, as validators take them.""" + +SchemaLeaf: TypeAlias = Callable[[object, "Schemas"], "JSONSchema | None"] +"""A caller's own shapes, written into a schema: asked first for every annotation, as a `Leaf` is, None to decline. + +Handed the schema being written, so a shape of the caller's own can hold +others, written with `of`, or be written once, in `$defs`, with `defined`. +""" + +DIALECT: Final = "https://json-schema.org/draft/2020-12/schema" +"""The dialect every schema is written in: JSON Schema draft 2020-12, as pydantic and zod write theirs.""" + +_KEYWORDS: Final[Mapping[str, str]] = { + "gt": "exclusiveMinimum", + "ge": "minimum", + "lt": "exclusiveMaximum", + "le": "maximum", +} +"""JSON Schema's keyword for each bound.""" + +_STRICTER: Final[Mapping[str, Callable[[float, float], float]]] = { + "exclusiveMinimum": max, + "minimum": max, + "exclusiveMaximum": min, + "maximum": min, +} +"""Of two bounds a keyword says, the one a value in both keeps within.""" + + +def no_schema_leaf(annotation: object, schemas: Schemas) -> JSONSchema | None: + """The schema leaf of a caller with no shapes of its own.""" + return None + + +class Schemas: + """One JSON Schema being written, and the `$defs` it holds so far. + + `of` writes an annotation as the checker reads it, asking the leaf + first at every depth, as `parser_for` asks a `Leaf`. A TypedDict and a + type alias are each written once, in `$defs`, under its name -- or its + name and a number, when another holds that one -- and referred to + wherever they occur, so one that holds itself is a schema that refers + to itself. `document` is the whole schema. + """ + + __slots__ = ("_defs", "_leaf", "_names", "_uses") + + def __init__(self, leaf: SchemaLeaf = no_schema_leaf) -> None: + self._leaf = leaf + self._defs: dict[str, JSONSchema] = {} + self._names: dict[object, str] = {} + self._uses: dict[str, int] = {} + + def of(self, annotation: object) -> JSONSchema: + """`annotation` as the checker reads it: its type, with the bounds and notes `Annotated` carries. + + A bound is the keyword JSON Schema has for it -- `Ge(0)` is + `minimum` -- and a `Doc` is the `description`. A type's bounds and + the bounds its `NewType` holds are both kept, as the checker holds + a value to both: where the two say one keyword, the stricter. An + annotation the checker reads is written; any other is a + `TypeError`, which a caller that vetted it through `parser` never + meets. + """ + inner, metadata = strip_annotation(annotation) + schema = self._type(inner) + if len(metadata) == 0: + return schema + notes = [ + item.documentation + for item in _unpacked(metadata) + if isinstance(item, typing_extensions.Doc) + ] + if len(notes) != 0: + schema = {**schema, "description": "\n\n".join(notes)} + for name, bound in constraints_of(metadata).items(): + keyword = _KEYWORDS[name] + held = cast("int | float | None", schema.get(keyword)) + schema = {**schema, keyword: bound if held is None else _STRICTER[keyword](held, bound)} + return schema + + def object_of(self, typeddict: type) -> JSONSchema: + """`typeddict` written in place: its keys, those it requires, and what any other key may hold.""" + keys = typeddict_keys(typeddict) + schema: JSONSchema = {"type": "object"} + if len(keys.members) != 0: + schema["properties"] = { + key: self.of(annotation) for key, (annotation, _) in keys.members.items() + } + required: list[JSONValue] = [key for key in keys.members if key in keys.required] + if len(required) != 0: + schema["required"] = required + if keys.closed: + schema["additionalProperties"] = False + elif not keys.open: + extra = self.of(keys.extra_items) + if len(extra) != 0: + schema["additionalProperties"] = extra + return schema + + def defined(self, key: object, name: str, write: Callable[[], JSONSchema]) -> JSONSchema: + """A reference to the entry in `$defs` for `key`, which `write` writes the first time `key` is asked for. + + The entry is named `name`, or `name` and a number when another key + holds that name, and it is reserved before it is written, so a + schema that holds itself refers to itself. One `write` fails to + write is not left reserved -- a later reference to it would be to an + empty schema, which takes anything -- and nor is anything it wrote + or counted before it failed. + """ + name_held = self._names.get(key) + if name_held is None: + name_held, number = name, 1 + while name_held in self._defs: + number += 1 + name_held = f"{name}{number}" + # What a failed write leaves is put back as it was: the entries + # it wrote of what it holds, and the uses it counted, as well as + # its own name. + before = (dict(self._defs), dict(self._names), dict(self._uses)) + self._names[key] = name_held + self._defs[name_held] = {} + try: + self._defs[name_held] = write() + except BaseException: + self._defs, self._names, self._uses = before + raise + self._uses[name_held] = self._uses.get(name_held, 0) + 1 + return {"$ref": _pointer(name_held)} + + def document(self, root: JSONSchema) -> JSONSchema: + """The whole schema: its dialect, `root`, and the `$defs`, by name. + + `root` is written in place when it refers to an entry nothing else + refers to, as pydantic writes a model that does not hold itself. + What comes back shares nothing with what was written, nor one part + of it with another, so a caller may change it where it likes. + """ + defs = dict(self._defs) + target = next((name for name in defs if root == {"$ref": _pointer(name)}), None) + if target is not None and self._uses[target] == 1: + root = defs.pop(target) + whole: JSONSchema = {"$schema": DIALECT, **root} + if len(defs) != 0: + whole["$defs"] = {name: defs[name] for name in sorted(defs)} + return cast("JSONSchema", copied(whole)) + + def _type(self, inner: object) -> JSONSchema: + """The schema of a type, `Annotated` peeled from it.""" + found = self._leaf(inner, self) + if found is not None: + return found + if inner is int: + return {"type": "integer"} + if inner is float: + return {"type": "number"} + if inner is bool: + return {"type": "boolean"} + if inner is str: + return {"type": "string"} + if inner is None or inner is types.NoneType: + return {"type": "null"} + if inner is JSONValue: + return {} + origin = get_origin(inner) + if origin is Literal: + # Sorted, as `_literal` sorts them: `get_args` reports a + # `Literal`'s values in the order the first one built wrote them. + values: list[JSONValue] = sorted(get_args(inner), key=repr) + return {"const": values[0]} if len(values) == 1 else {"enum": values} + if is_union(inner): + return {"anyOf": [self.of(branch) for branch in get_args(inner)]} + if origin is tuple: + return self._tuple(get_args(inner)) + if origin in (Mapping, dict): + value = self.of(get_args(inner)[1]) + if len(value) == 0: + return {"type": "object"} + return {"type": "object", "additionalProperties": value} + if isinstance(inner, type) and is_typeddict(inner): + typeddict = inner + return self.defined(typeddict, typeddict.__name__, lambda: self.object_of(typeddict)) + if isinstance(inner, NewType): + return self.of(inner.__supertype__) + if is_alias(inner): + alias = inner + return self.defined(alias, alias.__name__, lambda: self.of(alias_value(alias))) + msg = f"{inner!r} is not a shape JSON takes" + raise TypeError(msg) + + def _tuple(self, arguments: tuple[object, ...]) -> JSONSchema: + if len(arguments) == 2 and arguments[1] is Ellipsis: + items = self.of(arguments[0]) + return {"type": "array"} if len(items) == 0 else {"type": "array", "items": items} + if len(arguments) == 0: + return {"type": "array", "maxItems": 0} + return { + "type": "array", + "prefixItems": [self.of(argument) for argument in arguments], + "items": False, + "minItems": len(arguments), + } + + +def _pointer(name: str) -> str: + """The reference to the entry in `$defs` named `name`: a JSON pointer, escaped as a URI fragment.""" + escaped = name.replace("~", "~0").replace("/", "~1") + return f"#/$defs/{urllib.parse.quote(escaped, safe='')}" + + +def json_schema(shape: type) -> JSONSchema: + """The JSON Schema of the JSON `check` finds no problem with as `shape`, a TypedDict. + + Draft 2020-12, as a JSON object: arrays as lists, and `$schema` + first. A TypedDict is an object of its keys, those it requires, and + what any other key may hold -- nothing, in a closed one; a bound is + the keyword JSON Schema has for it, `Ge(0)` a `minimum`; a `Doc` is + the `description`, which is all that says one, as zod writes only + what `.describe()` said: a docstring is written for Python's readers; + a `Literal` is its values; a union is `anyOf` its branches. A + TypedDict or type alias is written once in `$defs`, under its name, + and referred to wherever it occurs, but for `shape` itself, which is + written in place unless it holds itself. + + One difference is JSON Schema's own: it takes a number with no + fraction, `1.0`, for an integer, where `check` wants `1`. `TypeError` + for a `shape` that is not a TypedDict, or holds something no parser + reads, as `check` raises it. + """ + if not is_typeddict(shape): + msg = f"{shape!r} is not a TypedDict" + raise TypeError(msg) + _checker(shape) + schemas = Schemas() + return schemas.document(schemas.of(shape)) __all__ = [ + "DIALECT", "Branch", + "Constraints", + "JSONSchema", "Leaf", "Loc", "Parsed", "Parser", "Qualifier", + "SchemaLeaf", + "Schemas", "Tag", "TypedDictKeys", "alias_value", "any_of", "check", + "constrained", + "constraints_of", "describe", "fixed_tuple", "has_shape", "is_alias", "is_integer", "is_union", + "json_schema", "mapping_of", "no_leaf", + "no_schema_leaf", "object_of", "one_of", "parser", @@ -973,5 +1434,6 @@ def check( "shape_of", "strip_annotation", "typeddict_keys", + "unqualified", "unread_in", ] diff --git a/packages/zarr-metadata/src/zarr_metadata/model/__init__.py b/packages/zarr-metadata/src/zarr_metadata/model/__init__.py index a148392584..ed467b40cd 100644 --- a/packages/zarr-metadata/src/zarr_metadata/model/__init__.py +++ b/packages/zarr-metadata/src/zarr_metadata/model/__init__.py @@ -1,68 +1,99 @@ """In-memory models for Zarr metadata documents. -Models are frozen dataclasses that hold a canonical, semantically lossless -representation of the JSON documents. Validators check a document's JSON +A model is a metadata document, as written and refined, and the scope it +was read in; what it hands out is read-only. Validators check a document's JSON structure and, in a v3 document, read each extension point (codecs, chunk grids, data types, ...) through the definition that claims its name in a scope, `CORE_AND_EXTENSIONS` unless a `context` is passed, and judge the -fill value against the data type it names, the chunk grid against -the shape, and the codecs as a pipeline, each against the chunk it is -handed. Each document concept gets a -`validate_*` function returning every problem found (a tuple of -`ValidationProblem`, each with a machine-readable `kind`), an `is_*` type -guard, and a `parse_*` function that narrows or raises -`MetadataValidationError`. Model `from_json` / `from_key_value` constructors -raise `MetadataValidationError` for every ingestion failure, including -missing store keys and undecodable bytes, and the v3 ones take the same -`context`. +fill value against the data type it names, the chunk grid against the +shape, and the codecs as a pipeline, each against the chunk it is +handed. Each document concept gets a `validate_*` function returning +every problem found (a tuple of `ValidationProblem`, each with a +machine-readable `kind`), an `is_*` type guard, and a `parse_*` function +that narrows or raises `MetadataValidationError`; a v3 array or group +document also gets `read_array_metadata_v3`, `read_group_metadata_v3` or +`read_array_metadata_v2`, +one read that returns what it read, the problems, and the model when +there are none. A store another writer made, holding a known writer +bug, is read by `read_repaired_node_metadata_v3`, which undoes each one +with `repair_node_metadata_v3` before the strict read and says what it +changed. `node_metadata_json_schema_v3` writes what the v3 +validators read as a JSON Schema, but for the rules. Model `from_json` / +`from_key_value` constructors raise +`MetadataValidationError` for every ingestion failure, including missing +store keys and undecodable bytes, and the v3 ones take the same +`context`. A v3 model is its document and the scope it was read in: +`ZarrV3ArrayMetadata(document, context=None)` reads the document in the +scope and raises `MetadataValidationError` with every problem, so no +model is built invalid; `to_json` writes the document as written; +`update` reads new members in the model's own scope; `with_context` and +`refined_in` read the document in another; `to_key_value` writes a model +as it is. A group's `consolidated_metadata` takes node models as entries, +each accepted when its claims refine into the group's scope and refused +at its path otherwise. Every reader takes `context=None` for the default +scope. """ from zarr_metadata._json import ( MetadataValidationError, ProblemKind, ValidationProblem, - is_json, - parse_json, - validate_json, ) +from zarr_metadata._sentinel import UNSET from zarr_metadata.model._array import ( ZarrV2ArrayMetadata, - ZarrV2ArrayMetadataPartial, + ZarrV2ArrayMetadataUpdate, ZarrV3ArrayMetadata, - ZarrV3ArrayMetadataPartial, - ZarrV3MetadataField, - ZarrV3NamedConfig, + ZarrV3ArrayMetadataUpdate, + read_array_metadata_v2, + read_array_metadata_v3, ) from zarr_metadata.model._group import ( ZarrV2ConsolidatedMetadata, ZarrV2GroupMetadata, - ZarrV2GroupMetadataPartial, + ZarrV2GroupMetadataUpdate, + ZarrV2NodeMetadata, ZarrV3ConsolidatedMetadata, + ZarrV3ConsolidatedMetadataInput, ZarrV3GroupMetadata, - ZarrV3GroupMetadataPartial, + ZarrV3GroupMetadataReading, + ZarrV3GroupMetadataUpdate, + ZarrV3NodeMetadata, + ZarrV3NodeMetadataInput, + ZarrV3NodeMetadataReading, + ZarrV3UnknownNodeReading, + is_group_metadata_v3, + node_metadata_from_json_v3, + node_metadata_from_key_value_v3, + parse_group_metadata_v3, + read_group_metadata_v3, + read_node_metadata_v3, + validate_group_metadata_v3, + validate_node_metadata_v3, +) +from zarr_metadata.model._json_schema import node_metadata_json_schema_v3 +from zarr_metadata.model._repair import ( + Repair, + RepairKind, + ZarrV2RepairedConsolidatedMetadataReading, + ZarrV3RepairedNodeMetadataReading, + read_repaired_consolidated_metadata_v2, + read_repaired_node_metadata_v3, + repair_consolidated_metadata_v2, + repair_node_metadata_v3, ) -from zarr_metadata.model._sentinel import UNSET from zarr_metadata.model._validation import ( - ARRAY_METADATA_OPTIONAL_KEYS_V3, - ARRAY_METADATA_REQUIRED_KEYS_V2, - ARRAY_METADATA_REQUIRED_KEYS_V3, - ARRAY_METADATA_STANDARD_KEYS_V3, - GROUP_METADATA_OPTIONAL_KEYS_V3, - GROUP_METADATA_REQUIRED_KEYS_V2, - GROUP_METADATA_REQUIRED_KEYS_V3, - GROUP_METADATA_STANDARD_KEYS_V3, + ZarrV2ArrayMetadataReading, + ZarrV3ArrayMetadataReading, is_array_metadata_v2, is_array_metadata_v3, is_group_metadata_v2, - is_group_metadata_v3, parse_array_metadata_v2, parse_array_metadata_v3, parse_group_metadata_v2, - parse_group_metadata_v3, validate_array_metadata_v2, validate_array_metadata_v3, validate_group_metadata_v2, - validate_group_metadata_v3, ) # Store keys are facts about the on-disk specs, so they are defined in the @@ -89,6 +120,14 @@ parse_metadata_field_v3, validate_metadata_field_v3, ) +from zarr_metadata.v3._hierarchy import ( + is_node_name_v3, + is_node_path_v3, + parse_node_name_v3, + parse_node_path_v3, + validate_node_name_v3, + validate_node_path_v3, +) from zarr_metadata.v3.array import ( ZARR_V3_ARRAY_METADATA_STORE_KEY, ZarrV3ArrayMetadataStoreKey, @@ -100,14 +139,6 @@ ) __all__ = [ - "ARRAY_METADATA_OPTIONAL_KEYS_V3", - "ARRAY_METADATA_REQUIRED_KEYS_V2", - "ARRAY_METADATA_REQUIRED_KEYS_V3", - "ARRAY_METADATA_STANDARD_KEYS_V3", - "GROUP_METADATA_OPTIONAL_KEYS_V3", - "GROUP_METADATA_REQUIRED_KEYS_V2", - "GROUP_METADATA_REQUIRED_KEYS_V3", - "GROUP_METADATA_STANDARD_KEYS_V3", "UNSET", "ZARR_V2_ARRAY_METADATA_STORE_KEY", "ZARR_V2_ATTRIBUTES_STORE_KEY", @@ -118,41 +149,67 @@ "ZARR_V3_GROUP_METADATA_STORE_KEY", "MetadataValidationError", "ProblemKind", + "Repair", + "RepairKind", "ValidationProblem", "ZarrV2ArrayMetadata", - "ZarrV2ArrayMetadataPartial", + "ZarrV2ArrayMetadataReading", "ZarrV2ArrayMetadataStoreKey", + "ZarrV2ArrayMetadataUpdate", "ZarrV2AttributesStoreKey", "ZarrV2ConsolidatedMetadata", "ZarrV2ConsolidatedMetadataStoreKey", "ZarrV2GroupMetadata", - "ZarrV2GroupMetadataPartial", "ZarrV2GroupMetadataStoreKey", + "ZarrV2GroupMetadataUpdate", + "ZarrV2NodeMetadata", + "ZarrV2RepairedConsolidatedMetadataReading", "ZarrV3ArrayMetadata", - "ZarrV3ArrayMetadataPartial", + "ZarrV3ArrayMetadataReading", "ZarrV3ArrayMetadataStoreKey", + "ZarrV3ArrayMetadataUpdate", "ZarrV3ConsolidatedMetadata", + "ZarrV3ConsolidatedMetadataInput", "ZarrV3GroupMetadata", - "ZarrV3GroupMetadataPartial", + "ZarrV3GroupMetadataReading", "ZarrV3GroupMetadataStoreKey", - "ZarrV3MetadataField", - "ZarrV3NamedConfig", + "ZarrV3GroupMetadataUpdate", + "ZarrV3NodeMetadata", + "ZarrV3NodeMetadataInput", + "ZarrV3NodeMetadataReading", + "ZarrV3RepairedNodeMetadataReading", + "ZarrV3UnknownNodeReading", "is_array_metadata_v2", "is_array_metadata_v3", "is_group_metadata_v2", "is_group_metadata_v3", - "is_json", "is_metadata_field_v3", + "is_node_name_v3", + "is_node_path_v3", + "node_metadata_from_json_v3", + "node_metadata_from_key_value_v3", + "node_metadata_json_schema_v3", "parse_array_metadata_v2", "parse_array_metadata_v3", "parse_group_metadata_v2", "parse_group_metadata_v3", - "parse_json", "parse_metadata_field_v3", + "parse_node_name_v3", + "parse_node_path_v3", + "read_array_metadata_v2", + "read_array_metadata_v3", + "read_group_metadata_v3", + "read_node_metadata_v3", + "read_repaired_consolidated_metadata_v2", + "read_repaired_node_metadata_v3", + "repair_consolidated_metadata_v2", + "repair_node_metadata_v3", "validate_array_metadata_v2", "validate_array_metadata_v3", "validate_group_metadata_v2", "validate_group_metadata_v3", - "validate_json", "validate_metadata_field_v3", + "validate_node_metadata_v3", + "validate_node_name_v3", + "validate_node_path_v3", ] diff --git a/packages/zarr-metadata/src/zarr_metadata/model/_array.py b/packages/zarr-metadata/src/zarr_metadata/model/_array.py index 6d3979dba3..5350b10cdd 100644 --- a/packages/zarr-metadata/src/zarr_metadata/model/_array.py +++ b/packages/zarr-metadata/src/zarr_metadata/model/_array.py @@ -2,110 +2,89 @@ from __future__ import annotations -import copy import dataclasses from collections.abc import Mapping -from dataclasses import dataclass, field -from typing import TYPE_CHECKING, Literal, TypeAlias, cast +from types import MappingProxyType +from typing import TYPE_CHECKING, Any, Final, cast from typing_extensions import TypedDict, Unpack +from zarr_metadata._common import ( + JSONValue, +) from zarr_metadata._json import ( MetadataValidationError, ValidationProblem, + copied, + frozen, + is_object, + json_text, + refined_object, + with_input, ) -from zarr_metadata.model._sentinel import UNSET +from zarr_metadata._sentinel import UNSET +from zarr_metadata.model._keyed import Keyed from zarr_metadata.model._validation import ( - ARRAY_METADATA_STANDARD_KEYS_V3, + ArrayMembersV2, + ArrayMembersV3, StoreKey, + ZarrV2ArrayMetadataReading, + ZarrV3ArrayMetadataReading, + dimension_lengths, dump_store_json, load_store_json, - parse_array_metadata_v2, - parse_array_metadata_v3, + read_array_v2, + read_array_v3, +) +from zarr_metadata.v2.array import ( + ZARR_V2_ARRAY_METADATA_STORE_KEY, + ZarrV2ArrayDimensionSeparator, + ZarrV2ArrayOrder, + ZarrV2DataTypeMetadata, ) -from zarr_metadata.v2.array import ZARR_V2_ARRAY_METADATA_STORE_KEY from zarr_metadata.v2.attributes import ZARR_V2_ATTRIBUTES_STORE_KEY -from zarr_metadata.v3._common import parse_metadata_field_v3 -from zarr_metadata.v3._registry import CORE_AND_EXTENSIONS -from zarr_metadata.v3.array import ZARR_V3_ARRAY_METADATA_STORE_KEY +from zarr_metadata.v2.codec import ZarrV2CodecMetadata +from zarr_metadata.v2.definition import CORE_V2, resolve_dtype_v2 +from zarr_metadata.v3._common import ZarrV3MetadataFieldJSON +from zarr_metadata.v3._definition import ( + AcceptedField, + ChunkGridDefinition, + ChunkKeyEncodingDefinition, + CodecDefinition, + DataTypeDefinition, + StorageTransformerDefinition, + UnclaimedField, + field_key, + fill_value_problems, + held, + spelled_canonically, +) +from zarr_metadata.v3._registry import ( + CORE_AND_EXTENSIONS, + Context, + ZarrV2Context, + ZarrV3Context, + scoped, +) +from zarr_metadata.v3._scope import Claims, Conflict, ScopeConflictError, claim_key, claims_of +from zarr_metadata.v3._scope import refines as refines_field +from zarr_metadata.v3.array import ZARR_V3_ARRAY_METADATA_STORE_KEY, ZarrV3ExtensionField if TYPE_CHECKING: - from zarr_metadata._common import JSONValue, ZarrV3NamedConfigJSON - from zarr_metadata.v2.array import ( - ZarrV2ArrayDimensionSeparator, - ZarrV2ArrayMetadataJSON, - ZarrV2ArrayMetadataStoreKey, - ZarrV2ArrayOrder, - ZarrV2DataTypeMetadata, - ) + from collections.abc import Iterable, Sequence + + from zarr_metadata._typed_json import Loc + from zarr_metadata.v2._definition import ZarrV2CodecDefinition, ZarrV2DataTypeDefinition + from zarr_metadata.v2.array import ZarrV2ArrayMetadataJSON, ZarrV2ArrayMetadataStoreKey from zarr_metadata.v2.attributes import ZarrV2AttributesStoreKey - from zarr_metadata.v2.codec import ZarrV2CodecMetadata - from zarr_metadata.v3._common import ZarrV3MetadataFieldJSON - from zarr_metadata.v3._registry import Context + from zarr_metadata.v3._definition import ResolvedField from zarr_metadata.v3.array import ( ZarrV3ArrayMetadataJSON, + ZarrV3ArrayMetadataJSONPartial, ZarrV3ArrayMetadataStoreKey, - ZarrV3ExtensionField, ) -@dataclass(frozen=True, slots=True, kw_only=True) -class ZarrV3NamedConfig: - """A normalized v3 metadata field with its reader obligation. - - Bare names and missing configurations normalize to an empty configuration. - Bare names and missing `must_understand` members normalize to the spec's - implicit `True` value (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L1571-L1573). - """ - - name: str - configuration: dict[str, JSONValue] - must_understand: bool = True - - def to_json(self) -> ZarrV3MetadataFieldJSON: - if not self.configuration and self.must_understand: - return self.name - # `configuration` is ReadOnly, so it is set in the literal rather than - # assigned afterwards. to_json output shares no mutable state with the - # model. - out: ZarrV3NamedConfigJSON = ( - {"name": self.name, "configuration": copy.deepcopy(self.configuration)} - if self.configuration - else {"name": self.name} - ) - if not self.must_understand: - out["must_understand"] = False - return out - - @classmethod - def from_json(cls, data: object) -> ZarrV3NamedConfig: - field = parse_metadata_field_v3(data) - if isinstance(field, str): - return cls(name=field, configuration={}, must_understand=True) - # A read model shares no mutable state with what it read. - configuration = copy.deepcopy(dict(field.get("configuration", {}))) - return cls( - name=field["name"], - configuration=configuration, - must_understand=field.get("must_understand", True), - ) - - -ZarrV3MetadataField: TypeAlias = ZarrV3NamedConfig -"""The in-memory model of one field of a v3 metadata document. - -This is the role-named alias for annotation positions: model fields and -consumer signatures should say `ZarrV3MetadataField` (the logical meaning) -rather than `ZarrV3NamedConfig` (the serialized form the field currently -takes). Today every metadata field normalizes to a named configuration plus -its reader obligation, so the alias is exactly `ZarrV3NamedConfig`; if a future -spec revision adds a field form that cannot be normalized to those values, -this alias widens to a union and annotation sites do not change. Mirrors the -raw-layer split between `ZarrV3NamedConfigJSON` (shape) and -`ZarrV3MetadataFieldJSON` (field union). -""" - - def must_understand_subset( extra_fields: Mapping[str, ZarrV3ExtensionField], ) -> dict[str, ZarrV3ExtensionField]: @@ -126,192 +105,174 @@ def must_understand_subset( } -class ZarrV3ArrayMetadataPartial(TypedDict, total=False): - """ - Partial form of the constructor-settable fields of `ZarrV3ArrayMetadata`. - - Every key is optional and typed with the model's own (not serialized) - value types, so it describes valid keyword arguments to - `ZarrV3ArrayMetadata.update`. The `init=False` fields `zarr_format` and - `node_type` are intentionally excluded, since they cannot be passed to - `dataclasses.replace`. +class ZarrV3ArrayMetadataUpdate(TypedDict, total=False, extra_items=ZarrV3ExtensionField | UNSET): + """The members `ZarrV3ArrayMetadata.update` puts in place: each as a document writes it, or `UNSET` to leave out one a document may leave out. - Drift between this type and the model's settable fields is prevented by - `tests/model/test_array.py::test_partial_keys_match_settable_model_fields`. + Those are `dimension_names`, `attributes`, `storage_transformers`, and + a member the spec does not define. """ shape: tuple[int, ...] + data_type: ZarrV3MetadataFieldJSON + chunk_grid: ZarrV3MetadataFieldJSON + chunk_key_encoding: ZarrV3MetadataFieldJSON fill_value: JSONValue - data_type: ZarrV3MetadataField - chunk_grid: ZarrV3MetadataField - codecs: tuple[ZarrV3MetadataField, ...] - chunk_key_encoding: ZarrV3MetadataField + codecs: tuple[ZarrV3MetadataFieldJSON, ...] + attributes: Mapping[str, JSONValue] | UNSET + storage_transformers: tuple[ZarrV3MetadataFieldJSON, ...] | UNSET dimension_names: tuple[str | None, ...] | UNSET - attributes: dict[str, JSONValue] - storage_transformers: tuple[ZarrV3MetadataField, ...] - extra_fields: dict[str, ZarrV3ExtensionField] - - -@dataclass(frozen=True, slots=True, kw_only=True) -class ZarrV3ArrayMetadata: - """In-memory model of a v3 array metadata document. - - A canonical, semantically lossless representation of the `zarr.json` - content for an array. Extension points (`data_type`, `chunk_grid`, - `chunk_key_encoding`, `codecs`, `storage_transformers`) are held as - `ZarrV3MetadataField` values (currently `ZarrV3NamedConfig` name, - configuration, and obligation records). `from_json` and - `from_key_value` read each through the definition that claims its name - in a scope -- `CORE_AND_EXTENSIONS` unless a `context` is passed -- and - the model holds what they read as written; `fill_value` is held - verbatim in its JSON form. Equivalent extension - spellings normalize to shorthand strings when configuration is empty and - understanding is required. + + +class ZarrV3ArrayMetadata(Keyed): + """A v3 array document, and the scope it was read in. + + The model is the pair: `to_json` is the document as written, refined + -- arrays as tuples, string keys -- and `context` the scope. Every + typed member is a view of the reading the pair gives: `data_type`, + `chunk_grid`, `chunk_key_encoding`, each codec and storage transformer + as the scope read it, `AcceptedField` by the definition that claims its name or + `UnclaimedField`; `shape`, `fill_value`, `dimension_names`, `attributes` and + `extra_fields` as the read refined them. Built only by reading: the + constructor reads `document` in `context` and raises + `MetadataValidationError` with every problem, so no model is invalid. + Two models are equal when their documents mean the same in their + scopes, as its key says; the scope itself takes no part. `update` + reads new members in the model's own scope; `with_context` and + `refined_in` read the document in another. A model pickles as its + pair, when the definitions its scope holds do: ones whose functions + are defined at a module's top level. """ - zarr_format: Literal[3] = field(default=3, init=False) - node_type: Literal["array"] = field(default="array", init=False) - shape: tuple[int, ...] - fill_value: JSONValue - data_type: ZarrV3MetadataField - chunk_grid: ZarrV3MetadataField - codecs: tuple[ZarrV3MetadataField, ...] - chunk_key_encoding: ZarrV3MetadataField - dimension_names: tuple[str | None, ...] | UNSET - attributes: dict[str, JSONValue] - storage_transformers: tuple[ZarrV3MetadataField, ...] - extra_fields: dict[str, ZarrV3ExtensionField] + __slots__ = ("_claims", "_context", "_document", "_members", "_reading", "_shown") + + zarr_format: Final = 3 + node_type: Final = "array" + + @property + def claims(self) -> Claims: + """What the reading claimed of each name the document writes, keyed as the scope files it.""" + return self._claims + + def __init__(self, document: object, context: ZarrV3Context | None = None) -> None: + scope = scoped(context, CORE_AND_EXTENSIONS) + reading, members = read_array_v3(document, scope) + if members is None: + raise MetadataValidationError(reading.problems) + self._adopt(refined_object(document), scope, reading, members) @classmethod - def create_default(cls, **overrides: Unpack[ZarrV3ArrayMetadataPartial]) -> ZarrV3ArrayMetadata: - """ - Create a default (empty) v3 array metadata model, with optional overrides. - - The default is a structurally-valid scalar `uint8` array — the array - analog of `list()` returning `[]`. Any field can be overridden by keyword - (the same fields accepted by `update`). Overriding `shape` without - `chunk_grid` derives a consistent default grid: one regular chunk - covering the array (`chunk_shape` equal to `shape`, with a length of - 1 for a dimension of length 0, since a chunk length is at least 1: - https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-grids/regular-grid/index.rst#L40). + def _of( + cls, + document: dict[str, JSONValue], + context: Context, + reading: ZarrV3ArrayMetadataReading, + members: ArrayMembersV3, + ) -> ZarrV3ArrayMetadata: + """A model of a document a read found nothing wrong with, holding that reading: no second read. The readers of this package build models through this, the private use pyright reports.""" + model = object.__new__(cls) + model._adopt(document, context, reading, members) + return model + + def _adopt( + self, + document: dict[str, JSONValue], + context: Context, + reading: ZarrV3ArrayMetadataReading, + members: ArrayMembersV3, + ) -> None: + self._document = document + self._context = context + # The reading holds the model it built, as `read_array_metadata_v3` + # hands it back, however the model was built. + self._reading = dataclasses.replace(reading, metadata=self) + self._members = members + # What the model shows of its members, read-only at every level. + self._shown = ( + frozen(members.fill_value), + frozen(members.attributes), + frozen(members.extra_fields), + ) + self._key = self._key_of() + self._claims = MappingProxyType(claims_of(reading.fields())) + + # --- the pair --------------------------------------------------------- - The derivation is deliberately one-way. A user-supplied `chunk_grid` - is an extension point and is taken verbatim — deriving `shape` from - it would require interpreting the grid's configuration, which this - layer never does (and cannot do for unrecognized grid names). So - overriding `chunk_grid` without `shape` keeps the scalar default - `shape=()`, which a grid of another rank does not fit: consistency - between the two is the caller's responsibility, so pass them - together. So is a fill value for an overridden `data_type`: - the default `fill_value` is `0`, which a data type whose fill value - is not an integer -- `bool`, `string`, a complex or struct type -- - refuses, so pass the two together; and so are its codecs: the - default `bytes` codec has no `endian`, which a data type whose - values take several bytes needs. + @property + def context(self) -> Context: + """The scope the document was read in, which `update` reads new members in.""" + return self._context + + @property + def reading(self) -> ZarrV3ArrayMetadataReading: + """The document as the scope read it: each field, the pipeline, the chunk each codec is handed.""" + return self._reading + + def to_json(self) -> ZarrV3ArrayMetadataJSON: + """The document as written, refined, sharing nothing with the model.""" + return cast("ZarrV3ArrayMetadataJSON", copied(self._document)) + + def to_key_value( + self, *, indent: int | str | None = None + ) -> Mapping[ZarrV3ArrayMetadataStoreKey, bytes]: + """The document as a store holds it: JSON bytes at `zarr.json`, indented by `indent`. + + `NaN`, `Infinity` and `-Infinity` in `attributes` are written as + those bare tokens, as zarr-python writes them, which a strict JSON + parser refuses. """ - if "shape" in overrides and "chunk_grid" not in overrides: - chunk_shape = tuple(max(length, 1) for length in overrides["shape"]) - overrides["chunk_grid"] = ZarrV3NamedConfig( - name="regular", configuration={"chunk_shape": chunk_shape} - ) - default = cls( - shape=(), - fill_value=0, - data_type=ZarrV3NamedConfig(name="uint8", configuration={}), - chunk_grid=ZarrV3NamedConfig(name="regular", configuration={"chunk_shape": ()}), - codecs=(ZarrV3NamedConfig(name="bytes", configuration={}),), - chunk_key_encoding=ZarrV3NamedConfig(name="default", configuration={}), - dimension_names=UNSET, - attributes={}, - storage_transformers=(), - extra_fields={}, + return {ZARR_V3_ARRAY_METADATA_STORE_KEY: dump_store_json(self._document, indent=indent)} + + def __repr__(self) -> str: + return f"{type(self).__name__}({self._document!r}, context={self._context!r})" + + def _key_of(self) -> tuple[object, ...]: + """What `==` and `hash` compare of a v3 array model: what its document means, as `array_key_of` says.""" + return array_key_of(self._reading, self._members) + + def _plain_key( + self, data_type: AcceptedField[DataTypeDefinition[Any]] | UnclaimedField + ) -> tuple[object, ...]: + """What `refines` compares of a model other than its fields, the fill value spelled as `data_type` -- the more informed side's -- spells it.""" + members = self._members + return ( + members.shape, + json_text(spelled_canonically(data_type, members.fill_value)), + members.dimension_names, + json_text(members.attributes), + json_text(members.extra_fields), ) - return default.update(**overrides) - def update(self, **kwargs: Unpack[ZarrV3ArrayMetadataPartial]) -> ZarrV3ArrayMetadata: - """ - Return a new `ZarrV3ArrayMetadata` with the given fields updated. + def __reduce__(self) -> tuple[type[ZarrV3ArrayMetadata], tuple[object, Context]]: + # The pair, read again on load: a model's reading never disagrees + # with its document. + return type(self), (self._document, self._context) - Only the constructor-settable fields listed in - `ZarrV3ArrayMetadataPartial` can be updated; any attempt to update - other fields (including the fixed `zarr_format` / `node_type`) is - rejected at the type level. Each given field fully replaces its - previous value, including `extra_fields`. + # --- typed views ------------------------------------------------------ - This is useful for test fixtures that want to override a few fields of a - base template without having to re-specify the entire document. + @property + def shape(self) -> tuple[int, ...]: + """The array's shape.""" + return self._members.shape - No re-validation is performed (`update` is `dataclasses.replace`), so - a repair or edit can produce an invalid document; validity is checked - on `from_json`, not on field replacement. - """ - return dataclasses.replace(self, **kwargs) - - def __post_init__(self) -> None: - overlap = set(self.extra_fields.keys()).intersection(ARRAY_METADATA_STANDARD_KEYS_V3) - if overlap: - raise MetadataValidationError( - [ - ValidationProblem( - ("extra_fields",), - "Extra fields cannot overlap with standard Zarr V3 array metadata fields", - "invalid_value", - ) - ] - ) + @property + def fill_value(self) -> JSONValue: + """The fill value as written, read-only at every level.""" + return self._shown[0] - def to_json(self) -> ZarrV3ArrayMetadataJSON: - # to_json output shares no mutable state with the model: every value - # that can hold a mutable container is deep-copied. - out: ZarrV3ArrayMetadataJSON = { - "zarr_format": self.zarr_format, - "node_type": self.node_type, - "shape": self.shape, - "fill_value": copy.deepcopy(self.fill_value), - "data_type": self.data_type.to_json(), - "chunk_grid": self.chunk_grid.to_json(), - "codecs": tuple(codec.to_json() for codec in self.codecs), - "chunk_key_encoding": self.chunk_key_encoding.to_json(), - } - if self.dimension_names is not UNSET: - out["dimension_names"] = self.dimension_names - if len(self.attributes) > 0: - out["attributes"] = copy.deepcopy(self.attributes) - if len(self.storage_transformers) > 0: - out["storage_transformers"] = tuple( - transformer.to_json() for transformer in self.storage_transformers - ) - # Extra fields are the TypedDict's `extra_items` (PEP 728). Assign them - # by key rather than `out.update(**...)`: type checkers understand the - # indexed-write path against `extra_items`, but not the `update(**...)` - # overload. - for key, value in self.extra_fields.items(): - out[key] = copy.deepcopy(value) - return out + @property + def dimension_names(self) -> tuple[str | None, ...] | UNSET: + """The dimension names; `UNSET` when the document writes none.""" + return self._members.dimension_names - @classmethod - def from_json( - cls, data: object, *, context: Context = CORE_AND_EXTENSIONS - ) -> ZarrV3ArrayMetadata: - # A read model shares no mutable state with what it read. - parsed = copy.deepcopy(parse_array_metadata_v3(data, context=context)) - extra_fields: dict[str, ZarrV3ExtensionField] = { - k: v for k, v in parsed.items() if k not in ARRAY_METADATA_STANDARD_KEYS_V3 - } - return cls( - shape=parsed["shape"], - fill_value=parsed["fill_value"], - data_type=ZarrV3NamedConfig.from_json(parsed["data_type"]), - chunk_grid=ZarrV3NamedConfig.from_json(parsed["chunk_grid"]), - codecs=tuple(ZarrV3NamedConfig.from_json(c) for c in parsed["codecs"]), - chunk_key_encoding=ZarrV3NamedConfig.from_json(parsed["chunk_key_encoding"]), - dimension_names=parsed.get("dimension_names", UNSET), - attributes=dict(parsed.get("attributes", {})), - storage_transformers=tuple( - ZarrV3NamedConfig.from_json(t) for t in parsed.get("storage_transformers", ()) - ), - extra_fields=extra_fields, - ) + @property + def attributes(self) -> Mapping[str, JSONValue]: + """The attributes, read-only at every level; empty when the document writes none.""" + return self._shown[1] + + @property + def extra_fields(self) -> Mapping[str, ZarrV3ExtensionField]: + """Each member the spec does not define, by name, read-only at every level.""" + return self._shown[2] @property def must_understand_fields(self) -> dict[str, ZarrV3ExtensionField]: @@ -325,35 +286,256 @@ def must_understand_fields(self) -> dict[str, ZarrV3ExtensionField]: """ return must_understand_subset(self.extra_fields) + @property + def data_type(self) -> AcceptedField[DataTypeDefinition[Any]] | UnclaimedField: + """The data type, as the scope read it.""" + return held(self._reading.data_type) + + @property + def chunk_grid(self) -> AcceptedField[ChunkGridDefinition[Any]] | UnclaimedField: + """The chunk grid, as the scope read it.""" + return held(self._reading.chunk_grid) + + @property + def chunk_key_encoding(self) -> AcceptedField[ChunkKeyEncodingDefinition[Any]] | UnclaimedField: + """The chunk key encoding, as the scope read it.""" + return held(self._reading.chunk_key_encoding) + + @property + def codecs(self) -> tuple[AcceptedField[CodecDefinition[Any]] | UnclaimedField, ...]: + """The codecs, each as the scope read it, in pipeline order.""" + return tuple(held(stage.codec) for stage in self._reading.pipeline) + + @property + def storage_transformers( + self, + ) -> tuple[AcceptedField[StorageTransformerDefinition[Any]] | UnclaimedField, ...]: + """The storage transformers, each as the scope read it.""" + return tuple(held(entry) for entry in self._reading.storage_transformers) + + # --- changing --------------------------------------------------------- + + def update(self, **members: Unpack[ZarrV3ArrayMetadataUpdate]) -> ZarrV3ArrayMetadata: + """This model with `members`, JSON, in place of the document's, `UNSET` leaving one out, read in this model's own scope. + + `MetadataValidationError` when the document they make has a + problem, so members that go together are passed together: a + `shape` with a grid that fits it. + """ + document: dict[str, object] = {**self._document, **members} + for key, value in members.items(): + if value is UNSET: + del document[key] + return type(self)(document, context=self._context) + + def with_context(self, context: ZarrV3Context | None = None) -> ZarrV3ArrayMetadata: + """This document read in `context`, whatever that changes: a gain, a loss, a conflict. + + `MetadataValidationError` when the document has a problem there. + The reading is kept when `context` reads every claim identically. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + if scope.disagreements(self._claims).agrees: + return self._of(self._document, scope, self._reading, self._members) + return type(self)(self._document, context=scope) + + def refined_in(self, context: ZarrV3Context | None = None) -> ZarrV3ArrayMetadata: + """This document read in `context`, which may claim what this scope left unclaimed and contradict nothing. + + `ScopeConflictError` naming each name `context` reads by another + definition, or by none, where this scope read it by one -- a loss + of meaning is refused as a conflict is -- and where each sits in + the document. `MetadataValidationError` when a name `context` + claims refuses what was written under it: a gain can surface a + problem. `with_context` reads the document in any scope. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + found = scope.disagreements(self._claims) + if len(found.conflicts) != 0: + raise ScopeConflictError(located_conflicts(self._reading.fields(), found.conflicts)) + return self.with_context(scope) + + def refines(self, other: ZarrV3ArrayMetadata) -> bool: + """Whether this model holds everything `other` holds: each field refines its counterpart, as `refines` orders fields -- the fields a field holds with it -- and every other member is the same, the fill value as the more informed data type spells it; a fill value that data type refuses is no refinement.""" + if type(other) is not type(self): + return False + if len(self.codecs) != len(other.codecs) or len(self.storage_transformers) != len( + other.storage_transformers + ): + return False + pairs = ( + (self.data_type, other.data_type), + (self.chunk_grid, other.chunk_grid), + (self.chunk_key_encoding, other.chunk_key_encoding), + *zip(self.codecs, other.codecs, strict=True), + *zip(self.storage_transformers, other.storage_transformers, strict=True), + ) + if not all(refines_field(mine, theirs) for mine, theirs in pairs): + return False + if len(fill_value_problems(self.data_type, other._members.fill_value)) != 0: + return False + return self._plain_key(self.data_type) == other._plain_key(self.data_type) + + # --- constructors ----------------------------------------------------- + + @classmethod + def create_default( + cls, + *, + context: ZarrV3Context | None = None, + **overrides: Unpack[ZarrV3ArrayMetadataJSONPartial], + ) -> ZarrV3ArrayMetadata: + """A scalar `uint8` array, or the one `overrides`, members of its document, make of it, read in `context`. + + `MetadataValidationError` when the document they make has a + problem, so members that go together are passed together: a data + type with a fill value of it, a grid with the shape it fits. The + default codec is `bytes` with a little `endian`, which takes a data + type of any fixed size. Overriding `shape` without `chunk_grid` + derives a consistent default grid: one regular chunk covering the + array (`chunk_shape` equal to `shape`, with a length of 1 for a + dimension of length 0, since "Chunk sizes must be greater than + zero", + https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-grids/regular-grid/index.rst#L40). + """ + # The grid derives from a shape the read takes; one it refuses is + # reported by the read, and derives nothing. + lengths, _ = dimension_lengths(overrides, "shape") + document: dict[str, object] = { + "zarr_format": 3, + "node_type": "array", + "shape": (), + "fill_value": 0, + "data_type": "uint8", + "chunk_grid": { + "name": "regular", + "configuration": {"chunk_shape": tuple(max(length, 1) for length in lengths or ())}, + }, + "codecs": ({"name": "bytes", "configuration": {"endian": "little"}},), + "chunk_key_encoding": {"name": "default"}, + } + return cls({**document, **overrides}, context=context) + + @classmethod + def from_json( + cls, data: object, *, context: ZarrV3Context | None = None + ) -> ZarrV3ArrayMetadata: + """The model of `data`, a v3 array document read in `context`. + + `MetadataValidationError` with every problem the read finds. + `read_array_metadata_v3` gives the reading this model is built + from, and the problems of a document with some. + """ + return cls(data, context=context) + @classmethod def from_key_value( - cls, mapping: Mapping[StoreKey, bytes], *, context: Context = CORE_AND_EXTENSIONS + cls, mapping: Mapping[StoreKey, bytes], *, context: ZarrV3Context | None = None ) -> ZarrV3ArrayMetadata: - return cls.from_json( - load_store_json(mapping, ZARR_V3_ARRAY_METADATA_STORE_KEY), context=context + """The model of the array document at `zarr.json` in `mapping`, read in `context`. + + `MetadataValidationError` when the key is missing, its bytes are not + JSON, or the document is not valid. + """ + return cls(load_store_json(mapping, ZARR_V3_ARRAY_METADATA_STORE_KEY), context=context) + + +def located_conflicts( + fields: Iterable[tuple[Loc, ResolvedField[Any]]], conflicts: Sequence[Conflict] +) -> tuple[Conflict, ...]: + """Each of `conflicts`, found against a reading's claims, once for each place among `fields` the name it is about sits: located, as a problem is.""" + located: list[Conflict] = [] + placed = list(fields) + for conflict in conflicts: + places = [(loc, field) for loc, field in placed if claim_key(field) == conflict.key] + if len(places) == 0: + located.append(conflict) + located.extend( + dataclasses.replace(conflict, loc=loc, written=field.name) for loc, field in places ) + return tuple(located) - def to_key_value( - self, *, indent: int | str | None = None, context: Context = CORE_AND_EXTENSIONS - ) -> Mapping[ZarrV3ArrayMetadataStoreKey, bytes]: - # A model built by hand is not validated: its document is written only - # if it reads as `from_json` reads one in `context`, and every problem - # is raised. - document = parse_array_metadata_v3(self.to_json(), context=context) - return {ZARR_V3_ARRAY_METADATA_STORE_KEY: dump_store_json(document, indent=indent)} +def array_key_of( + reading: ZarrV3ArrayMetadataReading, members: ArrayMembersV3 +) -> tuple[object, ...]: + """What a v3 array document means, as `reading` read it and `members` refine it: what `==` and `hash` compare of its model, and what two listings of one node are compared by. -class ZarrV2ArrayMetadataPartial(TypedDict, total=False): + Each field by its `field_key`, the fill value in its canonical spelling + as JSON text when a definition in scope read the data type and as + written when none did, and every other member as it is, the JSON ones + as text. """ - Partial form of the constructor-settable fields of `ZarrV2ArrayMetadata`. + data_type = held(reading.data_type) + fill_value = ( + spelled_canonically(data_type, members.fill_value) + if isinstance(data_type, AcceptedField) + else members.fill_value + ) + return ( + members.shape, + json_text(fill_value), + field_key(data_type), + field_key(held(reading.chunk_grid)), + tuple(field_key(held(stage.codec)) for stage in reading.pipeline), + field_key(held(reading.chunk_key_encoding)), + members.dimension_names, + json_text(members.attributes), + tuple(field_key(held(entry)) for entry in reading.storage_transformers), + json_text(members.extra_fields), + ) + - Every key is optional and typed with the model's own value types, so it - describes valid keyword arguments to `ZarrV2ArrayMetadata.update` and - `create_default`. The `init=False` field `zarr_format` is intentionally - excluded, since it cannot be passed to `dataclasses.replace`. +def read_array_metadata_v3( + value: object, *, context: ZarrV3Context | None = None +) -> ZarrV3ArrayMetadataReading: + """`value`, a v3 array document, as `context` read it, whatever it holds. - Drift between this type and the model's settable fields is prevented by - `tests/model/test_array.py::test_v2_partial_keys_match_settable_model_fields`. + Everything a read finds, in one: each extension point as `context` + read it -- `AcceptedField` by the definition that claims its name, `UnclaimedField`, + or `RefusedField` -- the chunks the codecs are handed, each codec with the + chunk it is handed, every problem `validate_array_metadata_v3` finds, + and, when there is none, the document's model, holding the same + reading. A policy over the fields, the core spec's alone, say, is a + walk over its `fields()`. A value that is not an object holds no + field. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + reading, members = read_array_v3(value, scope) + if members is None: + return reading + document = refined_object(value) + model = ZarrV3ArrayMetadata._of(document, scope, reading, members) # pyright: ignore[reportPrivateUsage] + return model.reading + + +def read_array_metadata_v2( + value: object, *, context: ZarrV2Context | None = None +) -> ZarrV2ArrayMetadataReading: + """`value`, a v2 array document, as `context` read it, `CORE_V2` when none is given, whatever it holds. + + Everything a read finds, in one: the dtype, the compressor and each + filter as the scope read them -- `AcceptedField` by the definition that claims + the typestr or id, `UnclaimedField`, or `RefusedField` -- every problem + `validate_array_metadata_v2` finds, and, when there is none, the + document's model. + """ + scope = scoped(context, CORE_V2) + reading, members = read_array_v2(value, scope) + if members is None: + return reading + document = refined_object(value) + if "dimension_separator" not in document: + document = {**document, "dimension_separator": "."} + model = ZarrV2ArrayMetadata._of(document, scope, reading, members) # pyright: ignore[reportPrivateUsage] + return model.reading + + +class ZarrV2ArrayMetadataUpdate(TypedDict, total=False, extra_items=JSONValue | UNSET): + """The members `ZarrV2ArrayMetadata.update` puts in place: each as a document writes it, or `UNSET` to leave out one a document may leave out. + + Those are `attributes` (no `.zattrs`), `dimension_separator` (read as + `"."`), and a member the spec does not define. """ shape: tuple[int, ...] @@ -363,159 +545,361 @@ class ZarrV2ArrayMetadataPartial(TypedDict, total=False): order: ZarrV2ArrayOrder compressor: ZarrV2CodecMetadata | None filters: tuple[ZarrV2CodecMetadata, ...] | None - dimension_separator: ZarrV2ArrayDimensionSeparator - attributes: dict[str, JSONValue] | UNSET - - -@dataclass(frozen=True, slots=True, kw_only=True) -class ZarrV2ArrayMetadata: - """In-memory model of a v2 array metadata document. - - A canonical, lossless representation of the `.zarray` content plus the - sibling `.zattrs` attributes. `dtype`, `compressor`, and `filters` are - held in their raw JSON forms and are never interpreted; `fill_value` is - held verbatim in its JSON form. `attributes` is `UNSET` when no - `.zattrs` file (or merged `attributes` key) exists — distinct from an - explicit empty `.zattrs`, which is `{}` and round-trips as a file. One - spelling normalization: a `.zarray` that omits `dimension_separator` - means `"."` by the v2 convention, and the model holds and re-emits that - value explicitly. + dimension_separator: ZarrV2ArrayDimensionSeparator | UNSET + attributes: Mapping[str, JSONValue] | UNSET + + +class ZarrV2ArrayMetadata(Keyed): + """A v2 array document, and the scope it was read in. + + The pair, as the v3 models are: `to_json` is the merged document -- + the `.zarray` members, and `attributes` when a `.zattrs` holds them -- + as written, refined, with one spelling put in: a `.zarray` that omits + `dimension_separator` means `"."` by the v2 convention + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L81-L86), + which the model holds and writes. `dtype`, `compressor` and each + filter are views of the reading: `AcceptedField` by the definition in scope + that claims the typestr or id, or `UnclaimedField`. `attributes` is + `UNSET` when no `.zattrs` exists, distinct from an empty one. A + member the spec does not define is kept, in `extra_fields`. Built + only by reading: the constructor reads `document` in `context` and + raises `MetadataValidationError` with every problem, so no model is + invalid. Two models are equal when their documents mean the same in + their scopes, as its key says. `update` reads new members in + the model's own scope; `with_context` and `refined_in` read the + document in another. A model pickles as its pair. """ - zarr_format: Literal[2] = field(default=2, init=False) - shape: tuple[int, ...] - dtype: ZarrV2DataTypeMetadata - chunks: tuple[int, ...] - fill_value: JSONValue - order: ZarrV2ArrayOrder - compressor: ZarrV2CodecMetadata | None - filters: tuple[ZarrV2CodecMetadata, ...] | None - # "." is the v2 convention's default for an ABSENT dimension_separator key; - # from_json normalizes absence to it (a semantics-preserving spelling - # normalization, like the v3 bare-string metadata-field form). The value - # is never None: the document grammar has no null spelling for this field. - dimension_separator: ZarrV2ArrayDimensionSeparator = field(default=".") - attributes: dict[str, JSONValue] | UNSET - - def update(self, **kwargs: Unpack[ZarrV2ArrayMetadataPartial]) -> ZarrV2ArrayMetadata: - """ - Return a new `ZarrV2ArrayMetadata` with the given fields updated. + __slots__ = ("_claims", "_context", "_document", "_members", "_reading", "_shown") - Only the constructor-settable fields listed in - `ZarrV2ArrayMetadataPartial` can be updated; the fixed `zarr_format` is - rejected at the type level. Each given field fully replaces its previous - value. - """ - return dataclasses.replace(self, **kwargs) + zarr_format: Final = 2 + + @property + def claims(self) -> Claims: + """What the reading claimed of each typestr and codec id the document writes, keyed as the scope files them.""" + return self._claims + + def __init__(self, document: object, context: ZarrV2Context | None = None) -> None: + scope = scoped(context, CORE_V2) + reading, members = read_array_v2(document, scope) + if members is None: + raise MetadataValidationError(reading.problems) + held = refined_object(document) + if "dimension_separator" not in held: + held = {**held, "dimension_separator": "."} + self._adopt(held, scope, reading, members) @classmethod - def create_default(cls, **overrides: Unpack[ZarrV2ArrayMetadataPartial]) -> ZarrV2ArrayMetadata: - """ - Create a default (empty) v2 array metadata model, with optional overrides. - - The default is a structurally-valid scalar `uint8` (`"|u1"`) array — the - array analog of `list()` returning `[]`. Any field can be overridden by - keyword (the same fields accepted by `update`). Overriding `shape` - without `chunks` derives `chunks` equal to `shape` (one chunk covering - the array). - - The derivation is deliberately one-way, matching the v3 model: - overriding `chunks` without `shape` keeps the scalar default - `shape=()`, and consistency between the two is the caller's - responsibility. - """ - if "shape" in overrides and "chunks" not in overrides: - overrides["chunks"] = tuple(overrides["shape"]) - default = cls( - shape=(), - dtype="|u1", - chunks=(), - fill_value=0, - order="C", - compressor=None, - filters=None, - attributes=UNSET, + def _of( + cls, + document: dict[str, JSONValue], + context: Context, + reading: ZarrV2ArrayMetadataReading, + members: ArrayMembersV2, + ) -> ZarrV2ArrayMetadata: + """A model of a document a read found nothing wrong with, holding that reading: no second read. The readers of this package build models through this, the private use pyright reports.""" + model = object.__new__(cls) + model._adopt(document, context, reading, members) + return model + + def _adopt( + self, + document: dict[str, JSONValue], + context: Context, + reading: ZarrV2ArrayMetadataReading, + members: ArrayMembersV2, + ) -> None: + self._document = document + self._context = context + self._reading = dataclasses.replace(reading, metadata=self) + self._members = members + # What the model shows of its members, read-only at every level. + self._shown = ( + frozen(members.fill_value), + UNSET if members.attributes is UNSET else frozen(members.attributes), + frozen(members.extra_fields), ) - return default.update(**overrides) + self._key = self._key_of() + self._claims = MappingProxyType(claims_of(reading.fields())) + + # --- the pair --------------------------------------------------------- + + @property + def context(self) -> Context: + """The scope the document was read in, which `update` reads new members in.""" + return self._context + + @property + def reading(self) -> ZarrV2ArrayMetadataReading: + """The document as the scope read it: the dtype, the compressor, each filter.""" + return self._reading def to_json(self) -> ZarrV2ArrayMetadataJSON: - """Return the merged in-memory document form. + """The merged document as written, refined, sharing nothing with the model. - `attributes` is included when set (even empty). This is not the - on-disk `.zarray` content: a conforming `.zarray` must exclude - `attributes` (they live in the sibling `.zattrs` file). Use - `to_key_value` to produce the spec-conforming split for storage + `attributes` is included when set, even empty. This is not the + on-disk `.zarray`, which excludes them: `to_key_value` splits the + document as a store holds it (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L323-L330). """ - # to_json output shares no mutable state with the model: every value - # that can hold a mutable container is deep-copied. - out: ZarrV2ArrayMetadataJSON = { - "zarr_format": self.zarr_format, - "shape": self.shape, - "dtype": self.dtype, - "order": self.order, - "chunks": self.chunks, - "fill_value": copy.deepcopy(self.fill_value), - "dimension_separator": self.dimension_separator, - "compressor": copy.deepcopy(self.compressor), - "filters": copy.deepcopy(self.filters), + return cast("ZarrV2ArrayMetadataJSON", copied(self._document)) + + def to_key_value( + self, *, indent: int | str | None = None + ) -> Mapping[ZarrV2ArrayMetadataStoreKey | ZarrV2AttributesStoreKey, bytes]: + """The document as a store holds it: `.zarray` without the attributes, and `.zattrs` with them when they are set, even empty.""" + zarray = {key: value for key, value in self._document.items() if key != "attributes"} + out: dict[ZarrV2ArrayMetadataStoreKey | ZarrV2AttributesStoreKey, bytes] = { + ZARR_V2_ARRAY_METADATA_STORE_KEY: dump_store_json(zarray, indent=indent) } - if self.attributes is not UNSET: - out["attributes"] = copy.deepcopy(self.attributes) + if "attributes" in self._document: + out[ZARR_V2_ATTRIBUTES_STORE_KEY] = dump_store_json( + self._document["attributes"], indent=indent + ) return out - @classmethod - def from_json(cls, data: object) -> ZarrV2ArrayMetadata: - # A read model shares no mutable state with what it read. - parsed = copy.deepcopy(parse_array_metadata_v2(data)) - return cls( - shape=parsed["shape"], - dtype=parsed["dtype"], - chunks=parsed["chunks"], - fill_value=parsed["fill_value"], - order=parsed["order"], - compressor=parsed["compressor"], - filters=parsed["filters"], - dimension_separator=parsed.get("dimension_separator", "."), - attributes=(dict(parsed["attributes"]) if "attributes" in parsed else UNSET), + def __repr__(self) -> str: + return f"{type(self).__name__}({self._document!r}, context={self._context!r})" + + def _key_of(self) -> tuple[object, ...]: + """What `==` and `hash` compare of a v2 array model: what its document means. + + Each field by its `field_key`, the fill value in its canonical spelling + as JSON text when a definition in scope read the dtype, and every other + member as it is, the JSON ones as text; `attributes` as `UNSET` when + there is no `.zattrs`. + """ + members = self._members + return ( + members.shape, + members.chunks, + members.order, + members.dimension_separator, + self._fill_value_key(), + field_key(self.dtype), + None if self.compressor is None else field_key(self.compressor), + None if self.filters is None else tuple(field_key(entry) for entry in self.filters), + UNSET if members.attributes is UNSET else json_text(members.attributes), + json_text(members.extra_fields), ) + def _plain_key( + self, dtype: AcceptedField[ZarrV2DataTypeDefinition[Any]] | UnclaimedField + ) -> tuple[object, ...]: + """What `refines` compares of a model other than its fields, the fill value spelled as `dtype` -- the more informed side's -- spells it.""" + members = self._members + return ( + members.shape, + members.chunks, + members.order, + members.dimension_separator, + json_text(spelled_canonically(dtype, members.fill_value)), + UNSET if members.attributes is UNSET else json_text(members.attributes), + json_text(members.extra_fields), + ) + + def _fill_value_key(self) -> str: + """What `==` compares of the fill value: its canonical spelling as JSON text when a definition in scope read the dtype, and the fill value as written when none did.""" + fill_value = self._members.fill_value + if isinstance(self.dtype, AcceptedField): + return json_text(spelled_canonically(self.dtype, fill_value)) + return json_text(fill_value) + + def __reduce__(self) -> tuple[type[ZarrV2ArrayMetadata], tuple[object, Context]]: + return type(self), (self._document, self._context) + + # --- typed views ------------------------------------------------------ + + @property + def shape(self) -> tuple[int, ...]: + """The array's shape.""" + return self._members.shape + + @property + def chunks(self) -> tuple[int, ...]: + """The shape of each chunk.""" + return self._members.chunks + + @property + def fill_value(self) -> JSONValue: + """The fill value as written, refined; read-only at every level.""" + return self._shown[0] + + @property + def order(self) -> ZarrV2ArrayOrder: + """The in-chunk layout, `"C"` or `"F"`.""" + return self._members.order + + @property + def dimension_separator(self) -> ZarrV2ArrayDimensionSeparator: + """What joins the chunk indices in a key: `"."` when the document writes none.""" + return self._members.dimension_separator + + @property + def attributes(self) -> Mapping[str, JSONValue] | UNSET: + """The user attributes a `.zattrs` holds, read-only at every level; `UNSET` when there is no `.zattrs`.""" + return self._shown[1] + + @property + def extra_fields(self) -> Mapping[str, JSONValue]: + """Every member the spec does not define, as written, read-only at every level.""" + return self._shown[2] + + @property + def dtype(self) -> AcceptedField[ZarrV2DataTypeDefinition[Any]] | UnclaimedField: + """The dtype as the scope read it: by its family's definition, or unclaimed.""" + return held(self._reading.dtype) + + @property + def compressor(self) -> AcceptedField[ZarrV2CodecDefinition[Any]] | UnclaimedField | None: + """The compressor as the scope read it; None when written as `null`.""" + compressor = self._reading.compressor + return None if compressor is None else held(compressor) + + @property + def filters( + self, + ) -> tuple[AcceptedField[ZarrV2CodecDefinition[Any]] | UnclaimedField, ...] | None: + """The filters, each as the scope read it; None when written as `null`.""" + filters = self._reading.filters + if filters is None: + return None + if filters is UNSET: + msg = "expected filters a read found nothing wrong with, got UNSET" + raise TypeError(msg) + return tuple(held(entry) for entry in filters) + + # --- changing --------------------------------------------------------- + + def update(self, **members: Unpack[ZarrV2ArrayMetadataUpdate]) -> ZarrV2ArrayMetadata: + """This model with `members`, JSON, in place of the document's, `UNSET` leaving one out, read in this model's own scope. + + `MetadataValidationError` when the document they make has a + problem, so members that go together are passed together: a + `dtype` with a fill value of it. + """ + document: dict[str, object] = {**self._document, **members} + for key, value in members.items(): + if value is UNSET: + del document[key] + return type(self)(document, context=self._context) + + def with_context(self, context: ZarrV2Context | None = None) -> ZarrV2ArrayMetadata: + """This document read in `context`, whatever that changes: a gain, a loss, a conflict. + + `MetadataValidationError` when the document has a problem there. + The reading is kept when `context` reads every claim identically. + """ + scope = scoped(context, CORE_V2) + if scope.disagreements(self._claims).agrees: + return self._of(self._document, scope, self._reading, self._members) + return type(self)(self._document, context=scope) + + def refined_in(self, context: ZarrV2Context | None = None) -> ZarrV2ArrayMetadata: + """This document read in `context`, which may claim what this scope left unclaimed and contradict nothing. + + `ScopeConflictError` naming each typestr or id `context` reads by + another definition, or by none, where this scope read it by one, + and where each sits in the document. `MetadataValidationError` + when a definition `context` claims refuses what was written. + """ + scope = scoped(context, CORE_V2) + found = scope.disagreements(self._claims) + if len(found.conflicts) != 0: + raise ScopeConflictError(located_conflicts(self._reading.fields(), found.conflicts)) + return self.with_context(scope) + + def refines(self, other: ZarrV2ArrayMetadata) -> bool: + """Whether this model holds everything `other` holds: each field refines its counterpart, a `null` compressor or filters only a `null`, and every other member is the same, the fill value as the more informed dtype spells it; a fill value that dtype refuses is no refinement.""" + if type(other) is not type(self): + return False + if (self.compressor is None) != (other.compressor is None): + return False + if (self.filters is None) != (other.filters is None): + return False + mine = () if self.filters is None else self.filters + theirs = () if other.filters is None else other.filters + if len(mine) != len(theirs): + return False + pairs = [(self.dtype, other.dtype), *zip(mine, theirs, strict=True)] + if self.compressor is not None and other.compressor is not None: + pairs.append((self.compressor, other.compressor)) + if not all(refines_field(one, another) for one, another in pairs): + return False + if len(fill_value_problems(self.dtype, other._members.fill_value)) != 0: + return False + return self._plain_key(self.dtype) == other._plain_key(self.dtype) + + # --- constructors ----------------------------------------------------- + + @classmethod + def create_default( + cls, *, context: ZarrV2Context | None = None, **overrides: Unpack[ZarrV2ArrayMetadataUpdate] + ) -> ZarrV2ArrayMetadata: + """A scalar `|u1` array, or the one `overrides`, members of its document, make of it, read in `context`. + + `MetadataValidationError` when the document they make has a + problem. Overriding `shape` without `chunks` derives `chunks` + equal to `shape`, one chunk covering the array; overriding `chunks` + without `shape` keeps the scalar default shape, which chunks of + another rank do not fit. A dtype given without a fill value takes + `0` when its family takes it, and `null` otherwise, which every + family takes. + """ + document: dict[str, object] = { + "zarr_format": 2, + "shape": (), + "chunks": (), + "dtype": "|u1", + "fill_value": 0, + "order": "C", + "compressor": None, + "filters": None, + "dimension_separator": ".", + } + given: dict[str, object] = dict(overrides) + if "shape" in given and "chunks" not in given: + lengths, _ = dimension_lengths(given, "shape") + if lengths is not None: + given["chunks"] = lengths + if "dtype" in given and "fill_value" not in given: + dtype, _ = resolve_dtype_v2(given["dtype"], context) + if len(fill_value_problems(dtype, 0)) != 0: + given["fill_value"] = None + merged = {key: value for key, value in {**document, **given}.items() if value is not UNSET} + return cls(merged, context=context) + + @classmethod + def from_json( + cls, data: object, *, context: ZarrV2Context | None = None + ) -> ZarrV2ArrayMetadata: + """The model of `data`, a v2 array document with its attributes under `attributes`, read in `context`. + + `MetadataValidationError` with every problem the read finds. + `read_array_metadata_v2` gives the reading this model is built + from, and the problems of a document with some. + """ + return cls(data, context=context) + @classmethod - def from_key_value(cls, mapping: Mapping[StoreKey, bytes]) -> ZarrV2ArrayMetadata: + def from_key_value( + cls, mapping: Mapping[StoreKey, bytes], *, context: ZarrV2Context | None = None + ) -> ZarrV2ArrayMetadata: + """The model of the array at `.zarray` in `mapping`, with the attributes at `.zattrs` when there is one, read in `context`. + + `MetadataValidationError` when `.zarray` is missing, bytes are not + JSON, `.zarray` holds `attributes`, or the document is not valid. + """ zarray_raw = load_store_json(mapping, ZARR_V2_ARRAY_METADATA_STORE_KEY) - if not isinstance(zarray_raw, Mapping): - return cls.from_json(zarray_raw) - zarray = cast("Mapping[str, object]", zarray_raw) + if not is_object(zarray_raw): + return cls(zarray_raw, context=context) + zarray = zarray_raw if "attributes" in zarray: - raise MetadataValidationError( - [ - ValidationProblem( - ("attributes",), - "unexpected document member", - "invalid_value", - ) - ] + refused = ValidationProblem( + ("attributes",), "unexpected document member", "invalid_value" ) + raise MetadataValidationError(with_input((refused,), zarray)) if ZARR_V2_ATTRIBUTES_STORE_KEY in mapping: zattrs = load_store_json(mapping, ZARR_V2_ATTRIBUTES_STORE_KEY) - return cls.from_json({**zarray, "attributes": zattrs}) - return cls.from_json(zarray) - - def to_key_value( - self, *, indent: int | str | None = None - ) -> Mapping[ZarrV2ArrayMetadataStoreKey | ZarrV2AttributesStoreKey, bytes]: - # Attributes live only in the sibling `.zattrs` file; the `.zarray` - # document must exclude them. The `.zattrs` key is present exactly - # when attributes are set (even empty) — UNSET emits no file. A model - # built by hand is not validated: its document is written only if it - # reads as `from_json` reads one, and every problem is raised. - document = parse_array_metadata_v2(self.to_json()) - zarray = {k: v for k, v in document.items() if k != "attributes"} - out: dict[ZarrV2ArrayMetadataStoreKey | ZarrV2AttributesStoreKey, bytes] = { - ZARR_V2_ARRAY_METADATA_STORE_KEY: dump_store_json(zarray, indent=indent) - } - if "attributes" in document: - out[ZARR_V2_ATTRIBUTES_STORE_KEY] = dump_store_json( - document["attributes"], indent=indent - ) - return out + return cls({**zarray, "attributes": zattrs}, context=context) + return cls(zarray, context=context) diff --git a/packages/zarr-metadata/src/zarr_metadata/model/_group.py b/packages/zarr-metadata/src/zarr_metadata/model/_group.py index fa56f633aa..2be9088767 100644 --- a/packages/zarr-metadata/src/zarr_metadata/model/_group.py +++ b/packages/zarr-metadata/src/zarr_metadata/model/_group.py @@ -2,166 +2,298 @@ from __future__ import annotations -import copy import dataclasses -from collections.abc import Mapping -from dataclasses import dataclass, field -from typing import TYPE_CHECKING, Literal, cast +from collections.abc import Callable, Mapping +from dataclasses import dataclass +from types import MappingProxyType +from typing import TYPE_CHECKING, Any, Final, Literal, TypeAlias, TypeGuard, cast -from typing_extensions import TypedDict, Unpack +from typing_extensions import TypeAliasType, TypedDict, Unpack +from zarr_metadata._common import JSONValue from zarr_metadata._json import ( MetadataValidationError, ValidationProblem, arrays_to_tuples, + copied, + frozen, + is_canonical_json, + is_json_object, + is_object, + json_text, + nested_past_the_levels, + not_an_object, + object_at, + outside_of, refine_json, refine_user_data, + refined_object, + shown, + shown_key, + with_input, + within, ) +from zarr_metadata._json import prefixed as _prefix +from zarr_metadata._sentinel import UNSET from zarr_metadata.model._array import ( + ZarrV2ArrayMetadata, ZarrV3ArrayMetadata, + array_key_of, + located_conflicts, must_understand_subset, + read_array_metadata_v3, ) -from zarr_metadata.model._sentinel import UNSET +from zarr_metadata.model._keyed import Keyed from zarr_metadata.model._validation import ( + GROUP_METADATA_REQUIRED_KEYS_V3, GROUP_METADATA_STANDARD_KEYS_V3, + ArrayMembersV3, StoreKey, + ZarrV3ArrayMetadataReading, + attributes_of, + check_literal, dump_store_json, + is_canonical_array_metadata_v3, load_store_json, + members_past_the_levels, + missing_keys, + other_members, parse_group_metadata_v2, - parse_group_metadata_v3, - validate_consolidated_metadata_v3, + read_array_v2, + read_array_v3, + reading_of, + unexpected_keys, + validate_group_metadata_v2, ) +from zarr_metadata.v2.array import ZARR_V2_ARRAY_METADATA_STORE_KEY from zarr_metadata.v2.attributes import ZARR_V2_ATTRIBUTES_STORE_KEY from zarr_metadata.v2.consolidated import ZARR_V2_CONSOLIDATED_METADATA_STORE_KEY +from zarr_metadata.v2.definition import CORE_V2 from zarr_metadata.v2.group import ZARR_V2_GROUP_METADATA_STORE_KEY -from zarr_metadata.v3._registry import CORE_AND_EXTENSIONS +from zarr_metadata.v3._hierarchy import NodeType, hierarchy_problems, path_faults, said +from zarr_metadata.v3._registry import ( + CORE_AND_EXTENSIONS, + Context, + ZarrV2Context, + ZarrV3Context, + scoped, +) +from zarr_metadata.v3._scope import ( + Claims, + Conflict, + ScopeConflictError, + claims_of, + definition_said, + kind_name, +) +from zarr_metadata.v3.array import ZarrV3ExtensionField from zarr_metadata.v3.consolidated import ZARR_V3_CONSOLIDATED_METADATA_KEY -from zarr_metadata.v3.group import ZARR_V3_GROUP_METADATA_STORE_KEY +from zarr_metadata.v3.group import ZARR_V3_GROUP_METADATA_STORE_KEY, ZarrV3GroupMetadataJSON if TYPE_CHECKING: - from zarr_metadata._common import JSONValue + from collections.abc import Iterator + + from zarr_metadata._typed_json import Loc from zarr_metadata.v2.attributes import ZarrV2AttributesStoreKey from zarr_metadata.v2.consolidated import ZarrV2ConsolidatedMetadataStoreKey from zarr_metadata.v2.group import ZarrV2GroupMetadataJSON, ZarrV2GroupMetadataStoreKey - from zarr_metadata.v3._registry import Context - from zarr_metadata.v3.array import ZarrV3ExtensionField + from zarr_metadata.v3._definition import ResolvedField + from zarr_metadata.v3.array import ZarrV3ArrayMetadataJSON from zarr_metadata.v3.consolidated import ZarrV3ConsolidatedMetadataJSON - from zarr_metadata.v3.group import ZarrV3GroupMetadataJSON, ZarrV3GroupMetadataStoreKey + from zarr_metadata.v3.group import ZarrV3GroupMetadataJSONPartial, ZarrV3GroupMetadataStoreKey -class ZarrV3GroupMetadataPartial(TypedDict, total=False): - """ - Partial form of the constructor-settable fields of `ZarrV3GroupMetadata`. +ZarrV3NodeMetadataInput = TypeAliasType( + "ZarrV3NodeMetadataInput", + "ZarrV3ArrayMetadataJSON | ZarrV3GroupMetadataJSON | ZarrV3ArrayMetadata | ZarrV3GroupMetadata", +) +"""What consolidated metadata lists at a path when given to a constructor or `update`: a document, or a model of it.""" + - Every key is optional and typed with the model's own value types, so it - describes valid keyword arguments to `ZarrV3GroupMetadata.update` and - `create_default`. The `init=False` fields `zarr_format` and `node_type` - are intentionally excluded, since they cannot be passed to - `dataclasses.replace`. +class ZarrV3ConsolidatedMetadataInput(TypedDict, closed=True): + """The `consolidated_metadata` member as a constructor or `update` takes it: as a document writes it, each entry a document or a node model. - Drift between this type and the model's settable fields is prevented by - `tests/model/test_group.py::test_group_partial_keys_match_settable_model_fields`. + A node model is accepted when the group's scope reads every claim of + it identically, or claims what the model's scope left unclaimed -- it + is then read again there -- and refused, with a problem at its path, + where the two scopes read a name differently, or the group's scope + leaves it unclaimed. """ - attributes: dict[str, JSONValue] - consolidated_metadata: ZarrV3ConsolidatedMetadata | UNSET - extra_fields: dict[str, ZarrV3ExtensionField] + kind: Literal["inline"] + must_understand: Literal[False] + metadata: Mapping[str, ZarrV3NodeMetadataInput] -@dataclass(frozen=True, slots=True, kw_only=True) -class ZarrV3GroupMetadata: - """In-memory model of a v3 group metadata document. +class ZarrV3GroupMetadataUpdate(TypedDict, total=False, extra_items=ZarrV3ExtensionField | UNSET): + """The members `ZarrV3GroupMetadata.update` puts in place: each as a document writes it, or `UNSET` to leave it out. - A canonical, semantically lossless representation of the `zarr.json` - content for a group. The `consolidated_metadata` reference-implementation - convention is modeled as a typed field holding thin child models, each - array read in the scope the group is read in; every other unknown - top-level key lands in `extra_fields` verbatim. + `consolidated_metadata` is given as a document writes it, each entry a + document or a node model, as `ZarrV3ConsolidatedMetadataInput` says, + or as another group's `ZarrV3ConsolidatedMetadata`, whose models are + taken; left out, the document's are read again as part of the whole. """ - zarr_format: Literal[3] = field(default=3, init=False) - node_type: Literal["group"] = field(default="group", init=False) - attributes: dict[str, JSONValue] - consolidated_metadata: ZarrV3ConsolidatedMetadata | UNSET - extra_fields: dict[str, ZarrV3ExtensionField] - - def __post_init__(self) -> None: - reserved = GROUP_METADATA_STANDARD_KEYS_V3 | {ZARR_V3_CONSOLIDATED_METADATA_KEY} - if set(self.extra_fields.keys()).intersection(reserved): - raise MetadataValidationError( - [ - ValidationProblem( - ("extra_fields",), - "Extra fields cannot overlap with standard Zarr V3 group metadata fields", - "invalid_value", - ) - ] - ) + attributes: Mapping[str, JSONValue] | UNSET + consolidated_metadata: ZarrV3ConsolidatedMetadataInput | ZarrV3ConsolidatedMetadata | UNSET - @classmethod - def create_default(cls, **overrides: Unpack[ZarrV3GroupMetadataPartial]) -> ZarrV3GroupMetadata: - """ - Create a default (empty) v3 group metadata model, with optional overrides. - The default is a structurally-valid group with no attributes — the group - analog of `list()` returning `[]`. Any field can be overridden by keyword - (the same fields accepted by `update`). - """ - default = cls(attributes={}, consolidated_metadata=UNSET, extra_fields={}) - return default.update(**overrides) +class ZarrV3GroupMetadata(Keyed): + """A v3 group document, and the scope it was read in. - def update(self, **kwargs: Unpack[ZarrV3GroupMetadataPartial]) -> ZarrV3GroupMetadata: - """ - Return a new `ZarrV3GroupMetadata` with the given fields updated. + The model is the pair, as `ZarrV3ArrayMetadata` is: `to_json` is the + document as written, refined, and `context` the scope. `attributes` + and `extra_fields` are views of what the read refined. The + `consolidated_metadata` reference-implementation convention is a + `ZarrV3ConsolidatedMetadata` view of the same pair: each document it + holds is a model of this scope, built from this one read. Built only + by reading: the constructor reads `document` in `context` and raises + `MetadataValidationError` with every problem, a nested document's + located under `consolidated_metadata.metadata.`. + """ - Only the constructor-settable fields listed in - `ZarrV3GroupMetadataPartial` can be updated; the fixed `zarr_format` / - `node_type` are rejected at the type level. Each given field fully - replaces its previous value, including `extra_fields`. - """ - return dataclasses.replace(self, **kwargs) + __slots__ = ( + "_claims", + "_consolidated", + "_context", + "_document", + "_members", + "_reading", + "_shown", + ) - def to_json(self) -> ZarrV3GroupMetadataJSON: - # to_json output shares no mutable state with the model: every value - # that can hold a mutable container is deep-copied. - out: ZarrV3GroupMetadataJSON = { - "zarr_format": self.zarr_format, - "node_type": self.node_type, - } - if len(self.attributes) > 0: - out["attributes"] = copy.deepcopy(self.attributes) - if self.consolidated_metadata is not UNSET: - # Consolidated metadata is a known non-core top-level JSON field. - out[ZARR_V3_CONSOLIDATED_METADATA_KEY] = self.consolidated_metadata.to_json() - for key, value in self.extra_fields.items(): - out[key] = copy.deepcopy(value) - return out + zarr_format: Final = 3 + node_type: Final = "group" + + @property + def claims(self) -> Claims: + """What the reading claimed of each name the document and its consolidated documents write, keyed as the scope files it.""" + return self._claims + + def __init__(self, document: object, context: ZarrV3Context | None = None) -> None: + scope = scoped(context, CORE_AND_EXTENSIONS) + reading, members = read_group_v3(document, scope) + if members is None or len(reading.problems) != 0: + raise MetadataValidationError(reading.problems) + self._adopt(refined_object(documents_for(document)), scope, reading, members) @classmethod - def from_json( - cls, data: object, *, context: Context = CORE_AND_EXTENSIONS + def _of( + cls, + document: dict[str, JSONValue], + context: Context, + reading: ZarrV3GroupMetadataReading, + members: GroupMembersV3, ) -> ZarrV3GroupMetadata: - # A read model shares no mutable state with what it read. - parsed = copy.deepcopy(parse_group_metadata_v3(data, context=context)) - consolidated_raw: object = parsed.get(ZARR_V3_CONSOLIDATED_METADATA_KEY, UNSET) - consolidated: ZarrV3ConsolidatedMetadata | UNSET - if consolidated_raw is UNSET or consolidated_raw is None: - # consolidated_metadata: null was written by a historical - # zarr-python bug; it gets no model representation. It is read as - # absence and never written back — repaired, not preserved. - consolidated = UNSET + """A model of a document a read found nothing wrong with, holding that reading: no second read. The readers of this package build models through this, the private use pyright reports.""" + model = object.__new__(cls) + model._adopt(document, context, reading, members) + return model + + def _adopt( + self, + document: dict[str, JSONValue], + context: Context, + reading: ZarrV3GroupMetadataReading, + members: GroupMembersV3, + ) -> None: + self._document = document + self._context = context + self._members = members + if members.attributes is None: + msg = "a group model holds members a read found nothing wrong with" + raise TypeError(msg) + # What the model shows of its members, read-only at every level. + self._shown = ( + frozen(members.attributes), + frozen(members.extra_fields), + ) + held: Mapping[str, ZarrV3NodeMetadataReading] = reading.consolidated + if members.consolidated is UNSET: + self._consolidated: ZarrV3ConsolidatedMetadata | UNSET = UNSET else: - consolidated = ZarrV3ConsolidatedMetadata.from_json(consolidated_raw, context=context) - extra_fields: dict[str, ZarrV3ExtensionField] = { - k: v - for k, v in parsed.items() - if k not in GROUP_METADATA_STANDARD_KEYS_V3 and k != ZARR_V3_CONSOLIDATED_METADATA_KEY - } - return cls( - attributes=dict(parsed.get("attributes", {})), - consolidated_metadata=consolidated, - extra_fields=extra_fields, + member = object_at(document, ZARR_V3_CONSOLIDATED_METADATA_KEY) + documents = object_at(member, "metadata") + models = _nested_models(documents, context, reading.consolidated, members.consolidated) + # One model per document, of this scope: each nested reading + # holds the model the group holds. + held = MappingProxyType( + { + path: (models[path].reading if path in models else nested) + for path, nested in reading.consolidated.items() + } + ) + self._consolidated = ZarrV3ConsolidatedMetadata._of( # pyright: ignore[reportPrivateUsage] + member, context, models + ) + # The reading holds the model it built, however the model was built; + # what it holds of the nested documents is read-only, as the model is. + self._reading = dataclasses.replace( + reading, consolidated=MappingProxyType(dict(held)), metadata=self ) + self._key = self._key_of() + self._claims = MappingProxyType(claims_of(reading.fields())) + + # --- the pair --------------------------------------------------------- + + @property + def context(self) -> Context: + """The scope the document was read in, which `update` reads new members in.""" + return self._context + + @property + def reading(self) -> ZarrV3GroupMetadataReading: + """The document as the scope read it: each document its consolidated metadata holds, as read.""" + return self._reading + + def to_json(self) -> ZarrV3GroupMetadataJSON: + """The document as written, refined, sharing nothing with the model.""" + return cast("ZarrV3GroupMetadataJSON", copied(self._document)) + + def to_key_value( + self, *, indent: int | str | None = None + ) -> Mapping[ZarrV3GroupMetadataStoreKey, bytes]: + """The document as a store holds it: JSON bytes at `zarr.json`, indented by `indent`. + + `NaN`, `Infinity` and `-Infinity` in `attributes` are written as + those bare tokens, as zarr-python writes them, which a strict JSON + parser refuses. + """ + return {ZARR_V3_GROUP_METADATA_STORE_KEY: dump_store_json(self._document, indent=indent)} + + def __repr__(self) -> str: + return f"{type(self).__name__}({self._document!r}, context={self._context!r})" + + def _key_of(self) -> tuple[object, ...]: + """What `==` and `hash` compare of a v3 group model: its attributes and extra fields as JSON text, and what its consolidated metadata holds, by its key.""" + consolidated = self.consolidated_metadata + members = self._members + return ( + json_text(members.attributes), + UNSET if consolidated is UNSET else consolidated._key, + json_text(members.extra_fields), + ) + + def __reduce__(self) -> tuple[type[ZarrV3GroupMetadata], tuple[object, Context]]: + # The pair, read again on load. + return type(self), (self._document, self._context) + + # --- typed views ------------------------------------------------------ + + @property + def attributes(self) -> Mapping[str, JSONValue]: + """The attributes, read-only at every level; empty when the document writes none.""" + return self._shown[0] + + @property + def extra_fields(self) -> Mapping[str, ZarrV3ExtensionField]: + """Each member the spec does not define, `consolidated_metadata` apart, by name: read-only at every level.""" + return self._shown[1] + + @property + def consolidated_metadata(self) -> ZarrV3ConsolidatedMetadata | UNSET: + """The `consolidated_metadata` member as a model of this scope; `UNSET` when the document writes none.""" + return self._consolidated @property def must_understand_fields(self) -> dict[str, ZarrV3ExtensionField]: @@ -175,284 +307,1487 @@ def must_understand_fields(self) -> dict[str, ZarrV3ExtensionField]: """ return must_understand_subset(self.extra_fields) + # --- changing --------------------------------------------------------- + + def update(self, **members: Unpack[ZarrV3GroupMetadataUpdate]) -> ZarrV3GroupMetadata: + """This model with `members`, JSON, in place of the document's, `UNSET` leaving one out, read in this model's own scope. + + A `consolidated_metadata` given is read; left out, the document's + is read again as part of the whole. `MetadataValidationError` when + the document they make has a problem. + """ + document: dict[str, object] = {**self._document, **members} + for key, value in members.items(): + if value is UNSET: + del document[key] + return type(self)(document, context=self._context) + + def with_context(self, context: ZarrV3Context | None = None) -> ZarrV3GroupMetadata: + """This document read in `context`, whatever that changes; `MetadataValidationError` when it has a problem there. The reading is kept when `context` reads every claim identically.""" + scope = scoped(context, CORE_AND_EXTENSIONS) + if scope.disagreements(self._claims).agrees: + return self._of(self._document, scope, self._reading, self._members) + return type(self)(self._document, context=scope) + + def refined_in(self, context: ZarrV3Context | None = None) -> ZarrV3GroupMetadata: + """This document read in `context`, which may claim what this scope left unclaimed and contradict nothing. + + `ScopeConflictError` naming each name `context` reads by another + definition, or by none, and where each sits, in the documents the + consolidated metadata holds too; `MetadataValidationError` when a + name `context` claims refuses what was written under it. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + found = scope.disagreements(self._claims) + if len(found.conflicts) != 0: + raise ScopeConflictError(located_conflicts(self._reading.fields(), found.conflicts)) + return self.with_context(scope) + + def refines(self, other: ZarrV3GroupMetadata) -> bool: + """Whether this model holds everything `other` holds: the same attributes and extra fields, and consolidated metadata whose every document refines its counterpart.""" + if type(other) is not type(self): + return False + mine, theirs = self._members, other._members + if json_text(mine.attributes) != json_text(theirs.attributes): + return False + if json_text(mine.extra_fields) != json_text(theirs.extra_fields): + return False + mine, theirs = self._consolidated, other._consolidated + if mine is UNSET or theirs is UNSET: + return mine is UNSET and theirs is UNSET + return mine.refines(theirs) + + # --- constructors ----------------------------------------------------- + + @classmethod + def create_default( + cls, + *, + context: ZarrV3Context | None = None, + **members: Unpack[ZarrV3GroupMetadataJSONPartial], + ) -> ZarrV3GroupMetadata: + """A group with no attributes, or the one `members` of its document make of it, read in `context`; `MetadataValidationError` when its document has a problem.""" + return cls({"zarr_format": 3, "node_type": "group", **members}, context=context) + + @classmethod + def from_json( + cls, data: object, *, context: ZarrV3Context | None = None + ) -> ZarrV3GroupMetadata: + """The model of `data`, a v3 group document read in `context`, with each document its consolidated metadata holds. + + `MetadataValidationError` with every problem the read finds. A + `consolidated_metadata` of `null`, which zarr-python 3.0 and 3.1 + wrote, is a value the document wrote, and no object: a problem, + as the spec says an object; `read_repaired_node_metadata_v3` reads + such a store. A member the spec does not define is held in + `extra_fields`. + """ + return cls(data, context=context) + @classmethod def from_key_value( - cls, mapping: Mapping[StoreKey, bytes], *, context: Context = CORE_AND_EXTENSIONS + cls, mapping: Mapping[StoreKey, bytes], *, context: ZarrV3Context | None = None ) -> ZarrV3GroupMetadata: - return cls.from_json( - load_store_json(mapping, ZARR_V3_GROUP_METADATA_STORE_KEY), context=context - ) + """The model of the group document at `zarr.json` in `mapping`, read in `context`. - def to_key_value( - self, *, indent: int | str | None = None, context: Context = CORE_AND_EXTENSIONS - ) -> Mapping[ZarrV3GroupMetadataStoreKey, bytes]: - # A model built by hand is not validated: its document is written only - # if it reads as `from_json` reads one in `context`, and every problem - # is raised. - document = parse_group_metadata_v3(self.to_json(), context=context) - return {ZARR_V3_GROUP_METADATA_STORE_KEY: dump_store_json(document, indent=indent)} - - -@dataclass(frozen=True, slots=True, kw_only=True) -class ZarrV3ConsolidatedMetadata: - """In-memory model of v3 inline consolidated metadata. - - Models the reference-implementation convention where consolidated metadata - is embedded as an extension field on a group's `zarr.json`. Each entry in - `metadata` is a complete child document, held as a thin array or group - model. `must_understand` is typed permissively as `bool` to mirror the - document shape, but only `False` is valid; this is enforced at runtime. + `MetadataValidationError` when the key is missing, its bytes are not + JSON, or the document is not valid. + """ + return cls(load_store_json(mapping, ZARR_V3_GROUP_METADATA_STORE_KEY), context=context) + + +class ZarrV3ConsolidatedMetadata(Keyed): + """A group's inline `consolidated_metadata` member, and the scope it was read in. + + Models the reference-implementation convention where consolidated + metadata is embedded as an extension field on a group's `zarr.json`. + `metadata` maps each path to the model of the complete document there, + array or group, of this scope: a view of the group's pair when a group + holds it, built from the group's one read; or of its own pair, when + the member is read on its own. `kind` is `inline` and `must_understand` + `False`, by declaration. The documents and the group make the + hierarchy below the group, the group its root, each at its node's + path without the leading `/`: the node at `/a/b` at `a/b`. """ - kind: Literal["inline"] = field(default="inline", init=False) - must_understand: bool = False - metadata: dict[str, ZarrV3ArrayMetadata | ZarrV3GroupMetadata] + __slots__ = ("_context", "_document", "_metadata") - def __post_init__(self) -> None: - if self.must_understand is not False: - raise MetadataValidationError( - [ - ValidationProblem( - ("must_understand",), - f"Invalid value for 'must_understand'. Expected False. " - f"Got {self.must_understand!r}.", - "invalid_value", - ) - ] - ) + kind: Final = "inline" + must_understand: Final = False + + def __init__(self, member: object, context: ZarrV3Context | None = None) -> None: + scope = scoped(context, CORE_AND_EXTENSIONS) + # The member sits under a group's key wherever it is read, so the + # levels a reader walks are counted from there, as in the group. + readings, members, problems = _read_consolidated_v3( + member, scope, (ZARR_V3_CONSOLIDATED_METADATA_KEY,) + ) + if len(problems) != 0: + raise MetadataValidationError(problems) + document = refined_object(_member_documents_for(member)) + documents = object_at(document, "metadata") + self._adopt(document, scope, _nested_models(documents, scope, readings, members)) + + @classmethod + def _of( + cls, + document: dict[str, JSONValue], + context: Context, + metadata: dict[str, ZarrV3NodeMetadata], + ) -> ZarrV3ConsolidatedMetadata: + """The member of a group a read found nothing wrong with, holding the models that read built. The readers of this package build models through this, the private use pyright reports.""" + model = object.__new__(cls) + model._adopt(document, context, metadata) + return model + + def _adopt( + self, + document: dict[str, JSONValue], + context: Context, + metadata: dict[str, ZarrV3NodeMetadata], + ) -> None: + self._document = document + self._context = context + self._metadata = metadata + self._key = self._key_of() + # Hidden from the readings, which hold their own models; see `metadata`. + + @property + def context(self) -> Context: + """The scope the documents were read in.""" + return self._context + + @property + def metadata(self) -> Mapping[str, ZarrV3NodeMetadata]: + """The model of each document, by its path below the group: a read-only view.""" + return MappingProxyType(self._metadata) def to_json(self) -> ZarrV3ConsolidatedMetadataJSON: - # `must_understand` is emitted as the literal False: the field is typed - # permissively as `bool`, but `__post_init__` guarantees the value. - return { - "kind": self.kind, - "must_understand": False, - "metadata": {key: node.to_json() for key, node in self.metadata.items()}, - } + """The member as written, refined, sharing nothing with the model.""" + return cast("ZarrV3ConsolidatedMetadataJSON", copied(self._document)) + + def __repr__(self) -> str: + return f"{type(self).__name__}({self._document!r}, context={self._context!r})" + + def _key_of(self) -> tuple[object, ...]: + """What `==` and `hash` compare of consolidated metadata: each document's key, by its path, in path order.""" + return tuple( + (path, node._key) + for path, node in sorted(self.metadata.items(), key=lambda item: item[0]) + ) + + def __reduce__(self) -> tuple[type[ZarrV3ConsolidatedMetadata], tuple[object, Context]]: + return type(self), (self._document, self._context) + + def refines(self, other: ZarrV3ConsolidatedMetadata) -> bool: + """Whether every document this holds refines the one `other` holds at the same path, and neither holds a path the other does not; False of what is not consolidated metadata.""" + if type(other) is not type(self): + return False + if self._metadata.keys() != other._metadata.keys(): + return False + return all( + _node_refines(self._metadata[path], other._metadata[path]) for path in self._metadata + ) @classmethod def from_json( - cls, data: object, *, context: Context = CORE_AND_EXTENSIONS + cls, data: object, *, context: ZarrV3Context | None = None ) -> ZarrV3ConsolidatedMetadata: - normalized = arrays_to_tuples(data) - problems = validate_consolidated_metadata_v3(normalized, context=context) - if len(problems) != 0: - raise MetadataValidationError(problems) - env = cast("Mapping[str, object]", normalized) - entries: dict[str, ZarrV3ArrayMetadata | ZarrV3GroupMetadata] = {} - for key, entry in cast("Mapping[str, object]", env["metadata"]).items(): - node_type = cast("Mapping[str, object]", entry).get("node_type") - if node_type == "array": - entries[key] = ZarrV3ArrayMetadata.from_json(entry, context=context) + """The model of `data`, a group's `consolidated_metadata` member, each document read once in `context`, as the array or group its `node_type` says; `MetadataValidationError` with every problem found.""" + return cls(data, context=context) + + +def _node_refines(node: ZarrV3NodeMetadata, other: ZarrV3NodeMetadata) -> bool: + """Whether `node` refines `other`, as the models of one kind refine each other; models of two kinds do not.""" + if isinstance(node, ZarrV3ArrayMetadata): + return isinstance(other, ZarrV3ArrayMetadata) and node.refines(other) + return isinstance(other, ZarrV3GroupMetadata) and node.refines(other) + + +def _no_documents() -> Mapping[str, ZarrV3NodeMetadataReading]: + """What a group whose consolidated metadata holds none, or that has none, holds: nothing.""" + return {} + + +@dataclass(frozen=True, slots=True) +class ZarrV3GroupMetadataReading: + """A v3 group document as a scope read it, whatever it holds: each document its consolidated metadata holds, as read, every problem, and the model when there is none.""" + + consolidated: Mapping[str, ZarrV3NodeMetadataReading] = dataclasses.field( + default_factory=_no_documents + ) + """Each document its consolidated metadata holds, as `read_node_metadata_v3` reads one, by its path.""" + problems: tuple[ValidationProblem, ...] = () + """Every reason the document is not a valid one.""" + metadata: ZarrV3GroupMetadata | None = None + """The document's model when there is no problem; None otherwise.""" + + def fields(self) -> Iterator[tuple[Loc, ResolvedField[Any]]]: + """Each field of each document its consolidated metadata holds, as read, with where it sits in this document.""" + for path, reading in self.consolidated.items(): + for loc, node in reading.fields(): + yield (ZARR_V3_CONSOLIDATED_METADATA_KEY, "metadata", path, *loc), node + + def __reduce__(self) -> tuple[Callable[..., object], tuple[object, ...]]: + # A reading that holds its model pickles and copies as the model + # does, and comes back as that model's own reading, so one model + # per document still; one without is built again from the dict its + # read-only view views, which pickles where the view does not. + if self.metadata is not None: + return (reading_of, (self.metadata,)) + return (_group_reading, (dict(self.consolidated), self.problems, None)) + + +def _group_reading( + consolidated: dict[str, ZarrV3NodeMetadataReading], + problems: tuple[ValidationProblem, ...], + metadata: ZarrV3GroupMetadata | None, +) -> ZarrV3GroupMetadataReading: + """A group reading built again from what `ZarrV3GroupMetadataReading.__reduce__` gives, its documents held read-only.""" + return ZarrV3GroupMetadataReading(MappingProxyType(consolidated), problems, metadata) + + +@dataclass(frozen=True, slots=True) +class ZarrV3UnknownNodeReading: + """A v3 document of no node type the spec defines -- its `node_type` missing, or neither `"array"` nor `"group"` -- or not an object at all: nothing else of it is read but its `zarr_format`, as its problems say. + + So a document of another format says so: zarr-python 2's draft of v3 + wrote a root `zarr.json` whose `zarr_format` is a URL, and a v2 + document names format 2. + """ + + problems: tuple[ValidationProblem, ...] + """Why it is no node.""" + + @property + def metadata(self) -> None: + """Its model: none, since no node type says which model it is.""" + return None + + def fields(self) -> Iterator[tuple[Loc, ResolvedField[Any]]]: + """Its fields as read: none, since none of them is read.""" + return iter(()) + + +ZarrV3NodeMetadataReading = TypeAliasType( + "ZarrV3NodeMetadataReading", + "ZarrV3ArrayMetadataReading | ZarrV3GroupMetadataReading | ZarrV3UnknownNodeReading", +) +"""A v3 `zarr.json` as `read_node_metadata_v3` reads it: as the array or group its `node_type` says, or as neither.""" + + +def read_node_metadata_v3( + value: object, *, context: ZarrV3Context | None = None +) -> ZarrV3NodeMetadataReading: + """`value`, a v3 `zarr.json`, read in `context` as the node its `node_type` says it is. + + The node type is the tag of a union, as pydantic's discriminator and + zod's discriminated union read one: an array is read as + `read_array_metadata_v3` reads it, a group as `read_group_metadata_v3` + does, and a document that says neither, or is not an object, is + `ZarrV3UnknownNodeReading`, with the problems, its `zarr_format`'s + among them. So no caller reads `node_type` from JSON it has not read, + and a document of another format says it is not v3. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + node_type, problems = _node_type(value) + if node_type == "array": + return read_array_metadata_v3(value, context=scope) + if node_type == "group": + return read_group_metadata_v3(value, context=scope) + return ZarrV3UnknownNodeReading(problems) + + +ZarrV3NodeMetadata = TypeAliasType( + "ZarrV3NodeMetadata", "ZarrV3ArrayMetadata | ZarrV3GroupMetadata" +) +"""The model of a v3 `zarr.json`: an array's or a group's, as its `node_type` says.""" + + +def node_metadata_from_json_v3( + data: object, *, context: ZarrV3Context | None = None +) -> ZarrV3NodeMetadata: + """The model of `data`, a v3 `zarr.json` read in `context`, as the node its `node_type` says. + + What `ZarrV3ArrayMetadata.from_json` or `ZarrV3GroupMetadata.from_json` + gives, as pydantic's `TypeAdapter` validates a discriminated union. + `MetadataValidationError` with every problem `read_node_metadata_v3` + finds, a `node_type` that says neither among them. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + reading = read_node_metadata_v3(data, context=scope) + if reading.metadata is None: + raise MetadataValidationError(reading.problems) + return reading.metadata + + +def node_metadata_from_key_value_v3( + mapping: Mapping[StoreKey, bytes], *, context: ZarrV3Context | None = None +) -> ZarrV3NodeMetadata: + """The model of the document at `zarr.json` in `mapping`, read in `context` as the node its `node_type` says, as `node_metadata_from_json_v3` reads one. + + `MetadataValidationError` when the key is missing, its bytes are not + JSON, or the document is not a valid array or group. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + # An array's document and a group's are both at `zarr.json`. + document = load_store_json(mapping, ZARR_V3_GROUP_METADATA_STORE_KEY) + return node_metadata_from_json_v3(document, context=scope) + + +def validate_node_metadata_v3( + value: object, *, context: ZarrV3Context | None = None +) -> tuple[ValidationProblem, ...]: + """Every reason `value` is not a valid v3 `zarr.json`: those `validate_array_metadata_v3` or `validate_group_metadata_v3` finds in the node its `node_type` says it is, or why it says neither.""" + scope = scoped(context, CORE_AND_EXTENSIONS) + return _read_node_v3(value, scope)[0].problems + + +def _read_node_v3( + value: object, context: Context, at: tuple[str | int, ...] = () +) -> tuple[ZarrV3NodeMetadataReading, ArrayMembersV3 | GroupMembersV3 | None]: + """`value` read as `read_node_metadata_v3` reads it, without models, and its members refined. + + `at` is where it sits in the document handed in, so the levels a + reader walks are counted from that one's root: a document past them + is the problem `_refine` reports for a container there, and not read. + """ + past = nested_past_the_levels(value, at) + if past is not None: + return ZarrV3UnknownNodeReading(within((past,), at)), None + node_type, problems = _node_type(value) + if node_type == "array": + return read_array_v3(value, context, at=at) + if node_type == "group": + return read_group_v3(value, context, at=at) + return ZarrV3UnknownNodeReading(problems), None + + +_NODE_TYPES: Final = ("array", "group") +"""The node types the spec defines.""" + + +def _node_type(value: object) -> tuple[str | None, tuple[ValidationProblem, ...]]: + """The node type `value` says it is, one of `_NODE_TYPES`; None, with the problems, when it says none of them, or is not an object. + + A document that says none is judged by its `zarr_format` too, so one + of another format says it is not v3. + """ + if not is_object(value): + return None, not_an_object(value) + document = value + node_type = document.get("node_type") + if isinstance(node_type, str) and node_type in _NODE_TYPES: + return node_type, () + problems = [ + *missing_keys(frozenset({"zarr_format"}), document), + *check_literal(document, "zarr_format", 3), + ] + if "node_type" not in document: + problems.append(ValidationProblem(("node_type",), "missing required key", "missing_key")) + else: + problems.append(outside_of(("node_type",), node_type, _NODE_TYPES)) + return None, with_input(problems, document) + + +@dataclass(frozen=True, slots=True) +class GroupMembersV3: + """What a read refined of a v3 group document, as the models hold it: its own members, and those of each document its consolidated metadata holds that a model can be built of.""" + + attributes: dict[str, JSONValue] | None + """Its attributes; None when they have a problem.""" + extra_fields: dict[str, JSONValue] + """Each member the spec does not define that is JSON.""" + consolidated: Mapping[str, ArrayMembersV3 | GroupMembersV3] | UNSET + """Each array its consolidated metadata holds that has no problem, and each group, by path; UNSET when it holds none.""" + + +def read_group_metadata_v3( + value: object, *, context: ZarrV3Context | None = None +) -> ZarrV3GroupMetadataReading: + """`value`, a v3 group document, as `context` read it, whatever it holds. + + Everything a read finds, in one: each document its consolidated + metadata holds, read once, as `read_array_metadata_v3` and this read + one; every problem `validate_group_metadata_v3` finds; and, when there + is none, the group's model, whose consolidated metadata holds the + models of those documents, which their readings hold too. A value that + is not an object holds nothing. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + reading, members = read_group_v3(value, scope) + if members is None: + return reading + return _with_models(reading, members, documents_for(value), scope) + + +def read_group_v3( + value: object, context: Context, *, at: tuple[str | int, ...] = () +) -> tuple[ZarrV3GroupMetadataReading, GroupMembersV3 | None]: + """`value`, a v3 group document, as `context` read it, without models, and its members refined; None when it is not an object. + + Every value is JSON: a field object built by hand, in a document + `consolidated_metadata` holds, is not, and is refused as + `read_array_v3` refuses one it is not told to hold. `at` is where the + document sits in the one handed in, as `read_array_v3` takes it: each + document its consolidated metadata holds is read from where it sits, + so a chain of them is bounded by the levels a reader walks, counted + from the outermost root. + """ + if not is_object(value): + return ZarrV3GroupMetadataReading(problems=not_an_object(value)), None + doc = value + found: list[ValidationProblem] = list(missing_keys(GROUP_METADATA_REQUIRED_KEYS_V3, doc)) + past = members_past_the_levels(doc, at) + found.extend(past.values()) + whole, doc = doc, {key: item for key, item in doc.items() if key not in past} + extra_fields, others = other_members( + doc, + GROUP_METADATA_STANDARD_KEYS_V3, + additional_reserved_keys=frozenset({ZARR_V3_CONSOLIDATED_METADATA_KEY}), + at=at, + ) + found.extend(others) + found.extend(check_literal(doc, "zarr_format", 3)) + found.extend(check_literal(doc, "node_type", "group")) + attributes: dict[str, JSONValue] | None = {} + if "attributes" in doc: + attributes, problems = attributes_of(doc["attributes"], at) + found.extend(problems) + raw = doc.get(ZARR_V3_CONSOLIDATED_METADATA_KEY, UNSET) + consolidated: dict[str, ZarrV3NodeMetadataReading] = {} + held: Mapping[str, ArrayMembersV3 | GroupMembersV3] | UNSET = UNSET + if raw is not UNSET: + consolidated, held, inside = _read_consolidated_v3( + raw, context, (*at, ZARR_V3_CONSOLIDATED_METADATA_KEY) + ) + found.extend(_prefix(ZARR_V3_CONSOLIDATED_METADATA_KEY, inside)) + reading = ZarrV3GroupMetadataReading(MappingProxyType(consolidated), with_input(found, whole)) + return reading, GroupMembersV3(attributes, extra_fields, held) + + +_CONSOLIDATED_MEMBERS: Final = ("kind", "must_understand", "metadata") +"""The members of an inline `consolidated_metadata`, in the order the convention declares them.""" + +_CONSOLIDATED_ENVELOPE: Final[dict[str, object]] = {"kind": "inline", "must_understand": False} +"""What the member declares, by declaration.""" + + +def documents_for(value: object) -> object: + """`value`, a v3 group document, with each node model its consolidated metadata lists replaced by that model's document, and a `ZarrV3ConsolidatedMetadata` given as the member by the member it holds; `value` itself when it holds none. + + What a group's own document is built from, so a document built of + models is JSON as any other, each child written as the child wrote it. + """ + if not is_object(value): + return value + document = value + given = document.get(ZARR_V3_CONSOLIDATED_METADATA_KEY) + member = _member_documents_for(given) + if member is given: + return value + return {**document, ZARR_V3_CONSOLIDATED_METADATA_KEY: member} + + +def _member_documents_for(member: object) -> object: + """A `consolidated_metadata` member with each node model it lists replaced by its document; `member` itself when it lists none.""" + if isinstance(member, ZarrV3ConsolidatedMetadata): + return member.to_json() + if not is_object(member): + return member + entries = member.get("metadata") + if not is_object(entries): + return member + replaced: dict[object, object] = {} + changed = False + for path, entry in entries.items(): + if isinstance(entry, (ZarrV3ArrayMetadata, ZarrV3GroupMetadata)): + replaced[path] = entry.to_json() + changed = True + else: + # A document listed here may list models of its own. + replaced[path] = documents_for(entry) + changed = changed or replaced[path] is not entry + if not changed: + return member + return {**member, "metadata": replaced} + + +def _read_node_model( + entry: ZarrV3ArrayMetadata | ZarrV3GroupMetadata, context: Context, at: Loc +) -> tuple[ZarrV3NodeMetadataReading, ArrayMembersV3 | GroupMembersV3 | None]: + """`entry`, a node model given where a document is listed, as `context` takes it: its own reading when `context` reads every claim of it identically, a read of its document when `context` claims more, and problems where `context` reads a name otherwise, or by none. + + A model may join a group when its claims refine into the group's + scope. A gain reads the document again there, so a problem a newly + claimed definition finds is reported where it sits. + """ + document = entry.to_json() + found = context.disagreements(entry.claims) + if len(found.conflicts) != 0: + # Read again in the group's scope, so the reading is of that scope, + # as every reading the group holds is; the conflicts are its + # problems, each where the field sits. + reading, _ = _read_node_v3(document, context, at) + written = {loc: field.name for loc, field in entry.reading.fields()} + problems = tuple( + ValidationProblem( + conflict.loc if conflict.loc is not None else (), + _conflict_said( + conflict, None if conflict.loc is None else written.get(conflict.loc) + ), + "invalid_value", + ) + for conflict in located_conflicts(entry.reading.fields(), found.conflicts) + ) + # The conflicts first: the cause, before what the re-read finds of it. + return _with_problems(reading, (*problems, *reading.problems)), None + if found.agrees: + # The model's own reading, unless the document sits too deep + # where it is listed, which a read from there reports. + _, past = refine_user_data(document, at) + if len(past) != 0: + return _read_node_v3(document, context, at) + moved = entry.with_context(context) + return moved.reading, moved._members # pyright: ignore[reportPrivateUsage] + return _read_node_v3(document, context, at) + + +def _conflict_said(conflict: Conflict, written: str | None) -> str: + """What a conflict between a model entry's scope and the group's says: the kind and the name as the document writes it, what read it, and how the group's scope reads it -- by another definition, told apart from the model's where the two print alike, or by none.""" + kind, filed = conflict.key + name = filed if written is None else written + claimed = definition_said(conflict.claimed) + head = f"expected a document read in the group's scope, got a model that reads the {kind_name(kind)} {name!r}" + if conflict.found is None: + return f"{head} by {claimed}, which the group's scope leaves unclaimed" + found = definition_said(conflict.found) + if claimed == found: + return f"{head} by another definition than the one the group's scope reads it by, {found}" + return f"{head} by {claimed}, which the group's scope reads by {found}" + + +def _with_problems( + reading: ZarrV3NodeMetadataReading, problems: tuple[ValidationProblem, ...] +) -> ZarrV3NodeMetadataReading: + """`reading`, a read of a document in the group's scope, with `problems` as its problems and no model.""" + if isinstance(reading, (ZarrV3GroupMetadataReading, ZarrV3ArrayMetadataReading)): + return dataclasses.replace(reading, problems=problems, metadata=None) + return dataclasses.replace(reading, problems=problems) + + +def _read_consolidated_v3( + value: object, context: Context, at: tuple[str | int, ...] = () +) -> tuple[ + dict[str, ZarrV3NodeMetadataReading], + dict[str, ArrayMembersV3 | GroupMembersV3], + tuple[ValidationProblem, ...], +]: + """An inline `consolidated_metadata` member, as `context` read it: each document it holds, read once, as `read_node_metadata_v3` reads one, by its path; the members of each a model can be built of; and every problem, located in the member. + + `at` is where the member sits in the document handed in. Each container + the reader descends into -- the member, its `metadata`, each document + -- is judged where it sits, as `_refine` judges one, so a chain of + documents is bounded by the levels a reader walks. + """ + if isinstance(value, ZarrV3ConsolidatedMetadata): + # Another group's member, given whole: its models at their paths. + value = {**_CONSOLIDATED_ENVELOPE, "metadata": dict(value.metadata)} + past = nested_past_the_levels(value, at) + if past is not None: + return {}, {}, within((past,), at) + if not is_object(value): + return {}, {}, (ValidationProblem((), "expected an object", "invalid_type"),) + env = value + # Missing members are reported in the order the envelope declares them. + problems: list[ValidationProblem] = [ + ValidationProblem((key,), "missing required key", "missing_key") + for key in _CONSOLIDATED_MEMBERS + if key not in env + ] + problems.extend(unexpected_keys(frozenset(_CONSOLIDATED_MEMBERS), env)) + problems.extend(check_literal(env, "kind", "inline")) + problems.extend(check_literal(env, "must_understand", False)) + readings: dict[str, ZarrV3NodeMetadataReading] = {} + members: dict[str, ArrayMembersV3 | GroupMembersV3] = {} + node_types: dict[str, NodeType | None] = {} + entries = env.get("metadata") + past = nested_past_the_levels(entries, (*at, "metadata")) + if "metadata" in env and not is_object(entries): + problems.append(ValidationProblem(("metadata",), "expected an object", "invalid_type")) + elif is_object(entries) and past is not None: + problems.extend(within((past,), at)) + elif is_object(entries): + for key, entry in entries.items(): + if not isinstance(key, str): + problems.append( + ValidationProblem( + ("metadata",), f"non-string key {shown_key(key)}", "invalid_type" + ) + ) + continue + faults = _key_problems(key) + problems.extend(faults) + if isinstance(entry, (ZarrV3ArrayMetadata, ZarrV3GroupMetadata)): + readings[key], child = _read_node_model(entry, context, (*at, "metadata", key)) else: - entries[key] = ZarrV3GroupMetadata.from_json(entry, context=context) - return cls(metadata=entries) + readings[key], child = _read_node_v3(entry, context, (*at, "metadata", key)) + if child is not None: + members[key] = child + problems.extend(_prefix("metadata", _prefix(key, readings[key].problems))) + if len(faults) == 0: + node_types[key] = _node_type_of(readings[key]) + problems.extend(_hierarchy_problems(node_types)) + for key in node_types: + problems.extend( + _nested_listing_problems( + key, readings[key], members.get(key), readings, members, node_types + ) + ) + return readings, members, tuple(problems) -class ZarrV2GroupMetadataPartial(TypedDict, total=False): +def _key_problems(key: str) -> list[ValidationProblem]: + """What keeps `key` from being where consolidated metadata keeps a document, said in one problem at the key. + + Consolidated metadata holds the hierarchy below its group, the group its + root, and the reference implementation keeps the document of each node + at the node's path in that hierarchy without its leading `/`: the node + at `/a/b` at the key `a/b`. """ - Partial form of the constructor-settable fields of `ZarrV2GroupMetadata`. + faults = _below_faults(key) + if len(faults) == 0: + return [] + message = f"expected the path of a node below the group, got {shown(key)}, which {said(faults)}" + return [ValidationProblem(("metadata", key), message, "invalid_value")] + - Every key is optional and typed with the model's own value types, so it - describes valid keyword arguments to `ZarrV2GroupMetadata.update` and - `create_default`. The `init=False` field `zarr_format` is intentionally - excluded, since it cannot be passed to `dataclasses.replace`. +def _nested_listing_problems( + key: str, + listed: ZarrV3NodeMetadataReading, + listed_members: ArrayMembersV3 | GroupMembersV3 | None, + readings: Mapping[str, ZarrV3NodeMetadataReading], + members: Mapping[str, ArrayMembersV3 | GroupMembersV3], + node_types: Mapping[str, NodeType | None], +) -> list[ValidationProblem]: + """What is wrong with the own consolidated listing of `listed`, the group at `key`, against the group's flat listing, `readings` and their `node_types`: each problem at the nested entry. - Drift between this type and the model's settable fields is prevented by - `tests/model/test_group.py::test_group_partial_keys_match_settable_model_fields`. + The reference implementation lists every node below the group in the + group's own listing, flat, and gives each group it lists an empty + listing of its own. So a node a listed group lists is one the group + lists too, at the joined key, of the same node type, and the same + document, as each is read: one it lists alone would be dropped by the + reference reader, one it lists as another type contradicts the tree, + and one it lists otherwise would give a reader two answers for one + node. Only the listing's own entries are judged: what a group listed + there lists in turn is that group's own to judge, when its document is + read, and a document with a problem of its own is judged by that. """ + problems: list[ValidationProblem] = [] + for path, entry in _reading_listing(listed).items(): + if len(_below_faults(path)) != 0: + # Its own reader reports a key that is no node's path. + continue + joined = f"{key}/{path}" + here = ("metadata", key, ZARR_V3_CONSOLIDATED_METADATA_KEY, "metadata", path) + if joined not in node_types: + message = ( + f"expected a node the group lists, got {shown(f'/{joined}')}, which " + f"{shown(f'/{key}')} lists alone" + ) + problems.append(ValidationProblem(here, message, "invalid_value")) + continue + flat, nested = node_types[joined], _node_type_of(entry) + if flat is not None and nested is not None and flat != nested: + message = ( + f"expected {_an(flat)}, as the group lists {shown(f'/{joined}')}, got {_an(nested)}" + ) + problems.append(ValidationProblem(here, message, "invalid_value")) + continue + own = _own_listing_members(listed_members).get(path) + flat_members = members.get(joined) + if ( + own is not None + and flat_members is not None + and not _read_alike(readings[joined], flat_members, entry, own) + ): + message = ( + f"expected the document the group lists at {shown(f'/{joined}')}, got one " + "that reads otherwise: one node, one document" + ) + problems.append(ValidationProblem(here, message, "invalid_value")) + return problems + + +def _own_listing_members( + listed: ArrayMembersV3 | GroupMembersV3 | None, +) -> Mapping[str, ArrayMembersV3 | GroupMembersV3]: + """The members of each document a listed group's own listing holds that a model can be built of; none for an array, a group listing nothing, or a document with a problem.""" + if isinstance(listed, GroupMembersV3) and listed.consolidated is not UNSET: + return listed.consolidated + return {} + - attributes: dict[str, JSONValue] | UNSET +def _read_alike( + reading: ZarrV3NodeMetadataReading, + members: ArrayMembersV3 | GroupMembersV3, + other: ZarrV3NodeMetadataReading, + other_members: ArrayMembersV3 | GroupMembersV3, +) -> bool: + """Whether two documents of one node read the same: arrays by what their models compare, groups by their own members, each listing judged where it sits.""" + if isinstance(reading, ZarrV3ArrayMetadataReading) and isinstance( + other, ZarrV3ArrayMetadataReading + ): + if not isinstance(members, ArrayMembersV3) or not isinstance(other_members, ArrayMembersV3): + return True + return array_key_of(reading, members) == array_key_of(other, other_members) + if isinstance(members, GroupMembersV3) and isinstance(other_members, GroupMembersV3): + return (json_text(members.attributes), json_text(members.extra_fields)) == ( + json_text(other_members.attributes), + json_text(other_members.extra_fields), + ) + return True + + +def _an(node_type: NodeType) -> str: + return "an array" if node_type == "array" else "a group" -@dataclass(frozen=True, slots=True, kw_only=True) -class ZarrV2GroupMetadata: - """In-memory model of a v2 group metadata document. +def _reading_listing(reading: ZarrV3NodeMetadataReading) -> Mapping[str, ZarrV3NodeMetadataReading]: + """What a reading lists in its own consolidated metadata: nothing, for an array's or one of no node type.""" + if isinstance(reading, ZarrV3GroupMetadataReading): + return reading.consolidated + return {} - A canonical, lossless representation of the `.zgroup` content plus the - sibling `.zattrs` attributes, folded into a single in-memory value - (mirroring the merged `ZarrV2GroupMetadataJSON` document form). `attributes` is - `UNSET` when no `.zattrs` file (or merged `attributes` key) exists — - distinct from an explicit empty `.zattrs`, which is `{}` and round-trips - as a file. + +def _hierarchy_problems(node_types: Mapping[str, NodeType | None]) -> list[ValidationProblem]: + """What keeps the documents consolidated metadata keeps and its group from making a hierarchy, the group its root, as `hierarchy_problems` judges one: each at the key it is about. + + `node_types` gives the node type of each document by its key, None for + a document of no node type the spec defines, and holds only keys + `_key_problems` finds nothing wrong with. """ + nodes: dict[str, NodeType | None] = {"/": "group"} + nodes.update((f"/{key}", node_type) for key, node_type in node_types.items()) + return [ + ValidationProblem(("metadata", cast("str", found.loc[0])[1:]), found.message, found.kind) + for found in hierarchy_problems(nodes) + ] + + +def _node_type_of(reading: ZarrV3NodeMetadataReading) -> NodeType | None: + """The node type a document says it is, as its reading tells; None when it says none.""" + if isinstance(reading, ZarrV3ArrayMetadataReading): + return "array" + if isinstance(reading, ZarrV3GroupMetadataReading): + return "group" + return None + + +def _below_faults(path: str) -> list[str]: + """What keeps `path` from being the path of a node below a group, relative to the group, each said.""" + if path == "": + return ["is the group's own"] + if path.startswith("/"): + return ['starts with "/"'] + return path_faults(f"/{path}") + + +def _with_models( + reading: ZarrV3GroupMetadataReading, + members: GroupMembersV3, + value: object, + context: Context, + at: Loc = (), +) -> ZarrV3GroupMetadataReading: + """`reading`, holding the model of each document its consolidated metadata holds that has no problem, and its own when it has none: each built from `value`, the document this read read, in `context`. + + A document with a problem a reader walks past -- a key that is no + string, a value that is no JSON, a level past the cap -- refines to + nothing as a whole; each document its consolidated metadata holds is + then refined on its own, from where it sits in the document handed in + -- `at` is where this one sits -- so the ones without a problem still + have their models, and a listed group with a problem of its own holds + those of its own listing. + """ + if len(reading.problems) == 0: + document = refined_object(value) + model = ZarrV3GroupMetadata._of(document, context, reading, members) # pyright: ignore[reportPrivateUsage] + return model.reading + if members.consolidated is UNSET or not is_object(value): + return reading + member = value.get(ZARR_V3_CONSOLIDATED_METADATA_KEY) + if not is_object(member): + return reading + entries = member.get("metadata") + if not is_object(entries): + return reading + held_entries = entries + documents: dict[str, JSONValue] = {} + for path in members.consolidated: + if path not in held_entries: + continue + entry, problems = refine_user_data( + held_entries[path], (*at, ZARR_V3_CONSOLIDATED_METADATA_KEY, "metadata", path) + ) + if len(problems) == 0 and isinstance(entry, Mapping): + documents[path] = entry + models = _nested_models( + documents, + context, + reading.consolidated, + {path: child for path, child in members.consolidated.items() if path in documents}, + ) + held = dict(reading.consolidated) + for path, child in members.consolidated.items(): + nested = reading.consolidated[path] + if path in models: + held[path] = models[path].reading + elif ( + path in held_entries + and isinstance(nested, ZarrV3GroupMetadataReading) + and isinstance(child, GroupMembersV3) + ): + # A listed group with a problem of its own -- one the reader + # refused, or one it walked past -- still holds, in its + # reading, a model of each document in its own listing that + # has none. + held[path] = _with_models( + nested, + child, + held_entries[path], + context, + (*at, ZARR_V3_CONSOLIDATED_METADATA_KEY, "metadata", path), + ) + return dataclasses.replace(reading, consolidated=MappingProxyType(held)) + + +def _nested_models( + documents: Mapping[str, JSONValue], + context: Context, + readings: Mapping[str, ZarrV3NodeMetadataReading], + members: Mapping[str, ArrayMembersV3 | GroupMembersV3], +) -> dict[str, ZarrV3NodeMetadata]: + """The model of each document in `documents`, a `metadata` member's, a model can be built of, in `context`: the one its reading holds already when that is of `context`, else built from its reading and members, not read again.""" + models: dict[str, ZarrV3NodeMetadata] = {} + for path, child in members.items(): + reading = readings[path] + held = reading.metadata + if held is not None and held.context is context: + models[path] = held + continue + document = object_at(documents, path) + if isinstance(reading, ZarrV3ArrayMetadataReading): + models[path] = ZarrV3ArrayMetadata._of( # pyright: ignore[reportPrivateUsage] + document, context, reading, cast("ArrayMembersV3", child) + ) + elif isinstance(reading, ZarrV3GroupMetadataReading) and len(reading.problems) == 0: + models[path] = ZarrV3GroupMetadata._of( # pyright: ignore[reportPrivateUsage] + document, context, reading, cast("GroupMembersV3", child) + ) + return models + + +def validate_group_metadata_v3( + value: object, *, context: ZarrV3Context | None = None +) -> tuple[ValidationProblem, ...]: + """Return every reason `value` is not a valid v3 group document. + + Unknown top-level keys are allowed (they map to `extra_fields`); a + reader must understand each one that does not say `must_understand: + false`, which the model reports as `must_understand_fields`. A + `consolidated_metadata` member, if present, is validated too: its + envelope, and each document it holds by its path, each array read as + `validate_array_metadata_v3` reads one, in `context`. These are the + `problems` of `read_group_metadata_v3`, which holds what was read to + find them. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + return read_group_v3(value, scope)[0].problems + + +def is_group_metadata_v3( + value: object, *, context: ZarrV3Context | None = None +) -> TypeGuard[ZarrV3GroupMetadataJSON]: + """Whether `value` is a v3 group document `validate_group_metadata_v3` finds nothing wrong with, written with tuples.""" + scope = scoped(context, CORE_AND_EXTENSIONS) + return ( + is_canonical_json(value, finite=False) + and _is_canonical_group_metadata_v3(value) + and not validate_group_metadata_v3(value, context=scope) + ) + + +def _is_canonical_group_metadata_v3(value: object) -> bool: + """Whether `value`, a group document, is written with the containers `ZarrV3GroupMetadataJSON` declares: each document its consolidated metadata lists as its TypedDict has it, arrays as `is_canonical_array_metadata_v3` asks, groups so again.""" + if not is_object(value) or not isinstance(value, dict): + return False + member = value.get(ZARR_V3_CONSOLIDATED_METADATA_KEY) + if not is_object(member): + return True + entries = member.get("metadata") + if not is_object(entries): + return True + for entry in entries.values(): + if not is_object(entry): + continue + node_type = entry.get("node_type") + if node_type == "array" and not is_canonical_array_metadata_v3(entry): + return False + if node_type == "group" and not _is_canonical_group_metadata_v3(entry): + return False + return True + - zarr_format: Literal[2] = field(default=2, init=False) - attributes: dict[str, JSONValue] | UNSET +def parse_group_metadata_v3( + value: object, *, context: ZarrV3Context | None = None +) -> ZarrV3GroupMetadataJSON: + """Return `value` narrowed to `ZarrV3GroupMetadataJSON`, or raise `MetadataValidationError`.""" + scope = scoped(context, CORE_AND_EXTENSIONS) + problems = validate_group_metadata_v3(value, context=scope) + if len(problems) != 0: + raise MetadataValidationError(problems) + return cast("ZarrV3GroupMetadataJSON", arrays_to_tuples(documents_for(value))) + + +class ZarrV2GroupMetadataUpdate(TypedDict, total=False): + """The members `ZarrV2GroupMetadata.update` puts in place: `attributes` as a `.zattrs` writes them, or `UNSET` for no `.zattrs`.""" + + attributes: Mapping[str, JSONValue] | UNSET + + +class ZarrV2GroupMetadata(Keyed): + """A v2 group document, and the scope it was read in. + + The pair, as the v3 models are: `to_json` is the merged document -- + `.zgroup`, and `attributes` when a `.zattrs` holds them -- as written, + refined. `attributes` is `UNSET` when no `.zattrs` exists, distinct + from an empty one. A group holds no field a scope reads, so the scope + is held for uniformity: `update` reads new attributes in it, and no + other scope conflicts with the reading. Built only by reading: the + constructor raises `MetadataValidationError` with every problem. Two + groups are equal when their attributes are written alike, as + `group_key_v2` says. + """ + + __slots__ = ("_attributes", "_context", "_document", "_shown") + + _attributes: dict[str, JSONValue] | UNSET + _context: Context + _document: dict[str, JSONValue] + _shown: Mapping[str, JSONValue] | UNSET + + zarr_format: Final = 2 + + @property + def claims(self) -> Claims: + """What the reading claimed: nothing, since a group holds no field.""" + return MappingProxyType({}) + + def __init__(self, document: object, context: ZarrV2Context | None = None) -> None: + scope = scoped(context, CORE_V2) + parsed = parse_group_metadata_v2(document, context=scope) + self._adopt(refined_object(document), scope, _v2_attributes(parsed)) @classmethod - def create_default(cls, **overrides: Unpack[ZarrV2GroupMetadataPartial]) -> ZarrV2GroupMetadata: - """ - Create a default (empty) v2 group metadata model, with optional overrides. + def _of( + cls, + document: dict[str, JSONValue], + context: Context, + attributes: dict[str, JSONValue] | UNSET, + ) -> ZarrV2GroupMetadata: + """A model of a document a read found nothing wrong with: no second read. The readers of this package build models through this, the private use pyright reports.""" + model = object.__new__(cls) + model._adopt(document, context, attributes) + return model - The default is a structurally-valid group with no attributes — the group - analog of `list()` returning `[]`. Any field can be overridden by keyword - (the same fields accepted by `update`). - """ - default = cls(attributes=UNSET) - return default.update(**overrides) + def _adopt( + self, + document: dict[str, JSONValue], + context: Context, + attributes: dict[str, JSONValue] | UNSET, + ) -> None: + self._document = document + self._context = context + self._attributes = attributes + self._shown = UNSET if attributes is UNSET else frozen(attributes) + self._key = self._key_of() - def update(self, **kwargs: Unpack[ZarrV2GroupMetadataPartial]) -> ZarrV2GroupMetadata: - """ - Return a new `ZarrV2GroupMetadata` with the given fields updated. + @property + def context(self) -> Context: + """The scope the document was read in, which `update` reads new attributes in.""" + return self._context - Only the constructor-settable fields listed in - `ZarrV2GroupMetadataPartial` can be updated; the fixed `zarr_format` - is rejected at the type level. Each given field fully replaces its - previous value. - """ - return dataclasses.replace(self, **kwargs) + @property + def attributes(self) -> Mapping[str, JSONValue] | UNSET: + """The user attributes a `.zattrs` holds, read-only at every level; `UNSET` when there is no `.zattrs`.""" + return self._shown def to_json(self) -> ZarrV2GroupMetadataJSON: - """Return the merged in-memory document form. + """The merged document as written, refined, sharing nothing with the model. - `attributes` is included when set (even empty). This is not the - on-disk `.zgroup` content: a conforming `.zgroup` must exclude - `attributes` (they live in the sibling `.zattrs` file). Use - `to_key_value` to produce the spec-conforming split for storage + `attributes` is included when set, even empty. This is not the + on-disk `.zgroup`, which excludes them: `to_key_value` splits the + document as a store holds it (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L313; https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L323-L330). """ - # to_json output shares no mutable state with the model. - out: ZarrV2GroupMetadataJSON = {"zarr_format": self.zarr_format} - if self.attributes is not UNSET: - out["attributes"] = copy.deepcopy(self.attributes) + return cast("ZarrV2GroupMetadataJSON", copied(self._document)) + + def to_key_value( + self, *, indent: int | str | None = None + ) -> Mapping[ZarrV2GroupMetadataStoreKey | ZarrV2AttributesStoreKey, bytes]: + """The document as a store holds it: `.zgroup` without the attributes, and `.zattrs` with them when they are set, even empty.""" + zgroup = {key: value for key, value in self._document.items() if key != "attributes"} + out: dict[ZarrV2GroupMetadataStoreKey | ZarrV2AttributesStoreKey, bytes] = { + ZARR_V2_GROUP_METADATA_STORE_KEY: dump_store_json(zgroup, indent=indent) + } + if "attributes" in self._document: + out[ZARR_V2_ATTRIBUTES_STORE_KEY] = dump_store_json( + self._document["attributes"], indent=indent + ) return out + def __repr__(self) -> str: + return f"{type(self).__name__}({self._document!r}, context={self._context!r})" + + def _key_of(self) -> tuple[object, ...]: + """What `==` and `hash` compare of a v2 group model: its attributes as JSON text, or `UNSET` when there is no `.zattrs`.""" + attributes = self._attributes + return (UNSET if attributes is UNSET else json_text(attributes),) + + def __reduce__(self) -> tuple[type[ZarrV2GroupMetadata], tuple[object, Context]]: + return type(self), (self._document, self._context) + + def update(self, **members: Unpack[ZarrV2GroupMetadataUpdate]) -> ZarrV2GroupMetadata: + """This model with `attributes` in place of the document's, `UNSET` leaving them out, read in this model's own scope; `MetadataValidationError` when the document they make has a problem.""" + document: dict[str, object] = {**self._document, **members} + for key, value in members.items(): + if value is UNSET: + del document[key] + return type(self)(document, context=self._context) + + def with_context(self, context: ZarrV2Context | None = None) -> ZarrV2GroupMetadata: + """This document read in `context`: the same group, holding that scope.""" + scope = scoped(context, CORE_V2) + return self._of(self._document, scope, self._attributes) + + def refined_in(self, context: ZarrV2Context | None = None) -> ZarrV2GroupMetadata: + """This document read in `context`: a group holds no field, so no scope conflicts with its reading, and this is `with_context`.""" + return self.with_context(context) + + def refines(self, other: ZarrV2GroupMetadata) -> bool: + """Whether this group holds everything `other` holds: its attributes written alike; False of what is not a v2 group.""" + return type(other) is type(self) and self._key == other._key + @classmethod - def from_json(cls, data: object) -> ZarrV2GroupMetadata: - # A read model shares no mutable state with what it read. - parsed = copy.deepcopy(parse_group_metadata_v2(data)) - return cls(attributes=(dict(parsed["attributes"]) if "attributes" in parsed else UNSET)) + def create_default( + cls, *, context: ZarrV2Context | None = None, **overrides: Unpack[ZarrV2GroupMetadataUpdate] + ) -> ZarrV2GroupMetadata: + """A group with no `.zattrs`, or the one `overrides` make of it, read in `context`; `MetadataValidationError` when the document they make has a problem.""" + given = {key: value for key, value in overrides.items() if value is not UNSET} + return cls({"zarr_format": 2, **given}, context=context) @classmethod - def from_key_value(cls, mapping: Mapping[StoreKey, bytes]) -> ZarrV2GroupMetadata: + def from_json( + cls, data: object, *, context: ZarrV2Context | None = None + ) -> ZarrV2GroupMetadata: + """The model of `data`, a v2 group document with its attributes under `attributes`, read in `context`; `MetadataValidationError` with every problem.""" + return cls(data, context=context) + + @classmethod + def from_key_value( + cls, mapping: Mapping[StoreKey, bytes], *, context: ZarrV2Context | None = None + ) -> ZarrV2GroupMetadata: + """The model of the group at `.zgroup` in `mapping`, with the attributes at `.zattrs` when there is one, read in `context`. + + `MetadataValidationError` when `.zgroup` is missing, bytes are not + JSON, `.zgroup` holds `attributes`, or the document is not valid. + """ zgroup_raw = load_store_json(mapping, ZARR_V2_GROUP_METADATA_STORE_KEY) - if not isinstance(zgroup_raw, Mapping): - return cls.from_json(zgroup_raw) - zgroup = cast("Mapping[str, object]", zgroup_raw) + if not is_object(zgroup_raw): + return cls(zgroup_raw, context=context) + zgroup = zgroup_raw if "attributes" in zgroup: - raise MetadataValidationError( - [ - ValidationProblem( - ("attributes",), - "unexpected document member", - "invalid_value", - ) - ] + # A key `.zgroup` does not declare: its attributes are `.zattrs`. + refused = ValidationProblem( + ("attributes",), "unexpected key 'attributes'", "unknown_key" ) + raise MetadataValidationError(with_input((refused,), zgroup)) if ZARR_V2_ATTRIBUTES_STORE_KEY in mapping: zattrs = load_store_json(mapping, ZARR_V2_ATTRIBUTES_STORE_KEY) - return cls.from_json({**zgroup, "attributes": zattrs}) - return cls.from_json(zgroup) + return cls({**zgroup, "attributes": zattrs}, context=context) + return cls(zgroup, context=context) - def to_key_value( - self, *, indent: int | str | None = None - ) -> Mapping[ZarrV2GroupMetadataStoreKey | ZarrV2AttributesStoreKey, bytes]: - # Attributes live only in the sibling `.zattrs` file; the `.zgroup` - # document must exclude them. The `.zattrs` key is present exactly - # when attributes are set (even empty) — UNSET emits no file. A model - # built by hand is not validated: its document is written only if it - # reads as `from_json` reads one, and every problem is raised. - document = parse_group_metadata_v2(self.to_json()) - zgroup = {k: v for k, v in document.items() if k != "attributes"} - out: dict[ZarrV2GroupMetadataStoreKey | ZarrV2AttributesStoreKey, bytes] = { - ZARR_V2_GROUP_METADATA_STORE_KEY: dump_store_json(zgroup, indent=indent) - } - if "attributes" in document: - out[ZARR_V2_ATTRIBUTES_STORE_KEY] = dump_store_json( - document["attributes"], indent=indent - ) - return out +ZarrV2NodeMetadata: TypeAlias = "ZarrV2ArrayMetadata | ZarrV2GroupMetadata" +"""The model of one node a v2 `.zmetadata` document holds: an array, or a group.""" -@dataclass(frozen=True, slots=True, kw_only=True) -class ZarrV2ConsolidatedMetadata: - """In-memory model of a v2 `.zmetadata` document. - The `metadata` map holds the flat file-keyed entries (`"path/.zarray"`, - `"path/.zattrs"`, ...) verbatim, preserving the normalized JSON tree. - Entries are deliberately NOT merged into per-node models: which nodes had - a `.zattrs` file at all is information the canonical representation must - keep. Interpreting entries into node models is consumer work. +class ZarrV2ConsolidatedMetadata(Keyed): + """A v2 `.zmetadata` document, and the scope its nodes were read in. + + `metadata` holds the flat file-keyed entries (`"path/.zarray"`, + `"path/.zattrs"`, ...) as written, refined: which nodes had a + `.zattrs` at all is kept. `nodes` is each `.zarray` or `.zgroup` + entry, merged with its sibling `.zattrs`, as a model of this scope, + keyed by the node's path, `""` for the root; a `.zattrs` with no + sibling is kept and makes no node, and any other entry is JSON, kept. + Built only by reading: the constructor raises `MetadataValidationError` + with every problem, each located under its entry. Two documents are + equal when each node means the same and the other entries are written + alike, as its key says; `refines`, `with_context` and + `refined_in` go through the nodes. """ - zarr_consolidated_format: Literal[1] = field(default=1, init=False) - metadata: dict[str, JSONValue] + __slots__ = ("_context", "_document", "_nodes", "_shown") + + _context: Context + _document: dict[str, JSONValue] + _nodes: dict[str, ZarrV2NodeMetadata] + _shown: Mapping[str, JSONValue] + + zarr_consolidated_format: Final = 1 + + def __init__(self, document: object, context: ZarrV2Context | None = None) -> None: + scope = scoped(context, CORE_V2) + entries, nodes, problems = _read_consolidated_v2(document, scope) + if len(problems) != 0: + raise MetadataValidationError(problems) + self._adopt({"zarr_consolidated_format": 1, "metadata": entries}, scope, nodes) + + @classmethod + def _of( + cls, document: dict[str, JSONValue], context: Context, nodes: dict[str, ZarrV2NodeMetadata] + ) -> ZarrV2ConsolidatedMetadata: + """A model of a document a read found nothing wrong with, holding the nodes that read built. The readers of this package build models through this, the private use pyright reports.""" + model = object.__new__(cls) + model._adopt(document, context, nodes) + return model + + def _adopt( + self, document: dict[str, JSONValue], context: Context, nodes: dict[str, ZarrV2NodeMetadata] + ) -> None: + self._document = document + self._context = context + self._nodes = nodes + self._shown = frozen(object_at(document, "metadata")) + self._key = self._key_of() + + @property + def context(self) -> Context: + """The scope the nodes were read in.""" + return self._context + + @property + def metadata(self) -> Mapping[str, JSONValue]: + """The entries as written, refined, by store key; read-only at every level.""" + return self._shown + + @property + def nodes(self) -> Mapping[str, ZarrV2NodeMetadata]: + """The model of each node, by its path below the root, `""` for the root: a read-only view.""" + return MappingProxyType(self._nodes) def to_json(self) -> dict[str, JSONValue]: - # to_json output shares no mutable state with the model. + """The `.zmetadata` document as written, refined, sharing nothing with the model.""" + return cast("dict[str, JSONValue]", copied(self._document)) + + def to_key_value( + self, *, indent: int | str | None = None + ) -> Mapping[ZarrV2ConsolidatedMetadataStoreKey, bytes]: + """The document as a store holds it: JSON bytes at `.zmetadata`, indented by `indent`.""" return { - "zarr_consolidated_format": self.zarr_consolidated_format, - "metadata": copy.deepcopy(self.metadata), + ZARR_V2_CONSOLIDATED_METADATA_STORE_KEY: dump_store_json(self._document, indent=indent) } - @classmethod - def from_json(cls, data: object) -> ZarrV2ConsolidatedMetadata: - normalized = arrays_to_tuples(data) - if not isinstance(normalized, Mapping): - raise MetadataValidationError( - [ValidationProblem((), "expected a mapping", "invalid_type")] + def __repr__(self) -> str: + return f"{type(self).__name__}({self._document!r}, context={self._context!r})" + + def _key_of(self) -> tuple[object, ...]: + """What `==` and `hash` compare of v2 consolidated metadata: each node by its path and its own key, and every other entry as JSON text.""" + nodes = self._nodes + return ( + tuple(sorted((path, node._key) for path, node in nodes.items())), + self._other_entries_text(), + ) + + def _other_entries_text(self) -> str: + """The entries no node is read from -- an orphan `.zattrs`, any other key -- as JSON text: what `==` compares of them.""" + entries = object_at(self._document, "metadata") + consumed: set[str] = set() + for path, names in _entries_by_path(entries)[0].items(): + if path in self._nodes: + consumed.update(names.values()) + return json_text({key: value for key, value in entries.items() if key not in consumed}) + + def __reduce__(self) -> tuple[type[ZarrV2ConsolidatedMetadata], tuple[object, Context]]: + return type(self), (self._document, self._context) + + def with_context(self, context: ZarrV2Context | None = None) -> ZarrV2ConsolidatedMetadata: + """This document with every node read in `context`, whatever that changes; `MetadataValidationError` when a node has a problem there.""" + scope = scoped(context, CORE_V2) + return type(self)(self._document, context=scope) + + def refined_in(self, context: ZarrV2Context | None = None) -> ZarrV2ConsolidatedMetadata: + """This document with every node read in `context`, which may claim what this scope left unclaimed and contradict nothing. + + `ScopeConflictError` naming each conflict, located at the node's + entry; `MetadataValidationError` when a gain surfaces a problem. + """ + scope = scoped(context, CORE_V2) + conflicts: list[Conflict] = [] + entries = object_at(self._document, "metadata") + by_path, _ = _entries_by_path(entries) + for path, node in self._nodes.items(): + if not isinstance(node, ZarrV2ArrayMetadata): + continue + found = scope.disagreements(node.claims) + key = by_path[path][ZARR_V2_ARRAY_METADATA_STORE_KEY] + conflicts.extend( + dataclasses.replace( + conflict, + loc=("metadata", key, *(() if conflict.loc is None else conflict.loc)), + ) + for conflict in located_conflicts(node.reading.fields(), found.conflicts) ) - doc = cast("Mapping[object, object]", normalized) - problems: list[ValidationProblem] = [ - ValidationProblem((key,), "missing required key", "missing_key") - for key in ("zarr_consolidated_format", "metadata") - if key not in doc - ] - problems.extend( - ValidationProblem((key,), "unexpected document member", "invalid_value") - if isinstance(key, str) - else ValidationProblem((), f"non-string document key {key!r}", "invalid_type") - for key in doc - if key not in {"zarr_consolidated_format", "metadata"} + if len(conflicts) != 0: + raise ScopeConflictError(conflicts) + return self.with_context(scope) + + def refines(self, other: ZarrV2ConsolidatedMetadata) -> bool: + """Whether every node this holds refines the one `other` holds at the same path, neither holds a path the other does not, and the other entries are written alike; False of what is not v2 consolidated metadata.""" + if type(other) is not type(self): + return False + if self._nodes.keys() != other._nodes.keys(): + return False + if self._other_entries_text() != other._other_entries_text(): + return False + return all(_v2_node_refines(self._nodes[path], other._nodes[path]) for path in self._nodes) + + @classmethod + def from_json( + cls, data: object, *, context: ZarrV2Context | None = None + ) -> ZarrV2ConsolidatedMetadata: + """The model of `data`, a `.zmetadata` document, its nodes read in `context`; `MetadataValidationError` with every problem.""" + return cls(data, context=context) + + @classmethod + def from_key_value( + cls, mapping: Mapping[StoreKey, bytes], *, context: ZarrV2Context | None = None + ) -> ZarrV2ConsolidatedMetadata: + """The model of the document at `.zmetadata` in `mapping`, read in `context`. + + `MetadataValidationError` when the key is missing, its bytes are not + JSON, or the document is not valid. + """ + return cls( + load_store_json(mapping, ZARR_V2_CONSOLIDATED_METADATA_STORE_KEY), context=context ) - if "zarr_consolidated_format" in doc and ( - not isinstance(doc["zarr_consolidated_format"], int) - or isinstance(doc["zarr_consolidated_format"], bool) - or doc["zarr_consolidated_format"] != 1 - ): + + +def _v2_node_refines(node: ZarrV2NodeMetadata, other: ZarrV2NodeMetadata) -> bool: + """Whether `node` refines `other`, as models of one kind refine each other; models of two kinds do not.""" + if isinstance(node, ZarrV2ArrayMetadata): + return isinstance(other, ZarrV2ArrayMetadata) and node.refines(other) + return isinstance(other, ZarrV2GroupMetadata) and node.refines(other) + + +_NODE_FILES: Final = ( + ZARR_V2_ARRAY_METADATA_STORE_KEY, + ZARR_V2_GROUP_METADATA_STORE_KEY, + ZARR_V2_ATTRIBUTES_STORE_KEY, +) + + +def _entries_by_path( + entries: Mapping[str, JSONValue], +) -> tuple[dict[str, dict[str, str]], list[ValidationProblem]]: + """The `.zarray`, `.zgroup` and `.zattrs` entries, by node path, then by file: the key each sits under; and a problem for each second key naming one file of one node, which would otherwise go unread.""" + by_path: dict[str, dict[str, str]] = {} + problems: list[ValidationProblem] = [] + for key in entries: + path, _, name = key.rpartition("/") + if name not in _NODE_FILES: + continue + # A node's path, as a store names it: segments joined by one `/`. + segments = [segment for segment in path.split("/") if segment != ""] + if any(segment in (".", "..") for segment in segments): problems.append( ValidationProblem( - ("zarr_consolidated_format",), - f"expected 1, got {doc['zarr_consolidated_format']!r}", + ("metadata", key), + f"expected the path of a node, got {key!r}, which has a '.' or '..' segment", "invalid_value", ) ) - refined: dict[str, JSONValue] = {} - if "metadata" in doc: - entries = doc["metadata"] - if not isinstance(entries, Mapping) or not all( - isinstance(k, str) for k in cast("Mapping[object, object]", entries) - ): - problems.append( - ValidationProblem( - ("metadata",), "expected a mapping with string keys", "invalid_type" - ) + continue + path = "/".join(segments) + files = by_path.setdefault(path, {}) + if name in files: + problems.append( + ValidationProblem( + ("metadata", key), + f"a second {name} for the node at {path!r}, which {files[name]!r} is", + "invalid_value", + ) + ) + continue + files[name] = key + return by_path, problems + + +def _read_consolidated_v2( + data: object, context: Context +) -> tuple[dict[str, JSONValue], dict[str, ZarrV2NodeMetadata], tuple[ValidationProblem, ...]]: + """`data`, a `.zmetadata` document: each entry as read, the model of each node read in `context`, and every problem, located in the document. + + Each entry is the document its key names: a `.zattrs` is user data, + any other is JSON by RFC 8259; a `.zarray` or `.zgroup` is read as the + document it is, merged with its sibling `.zattrs`, its problems under + its entry and the attributes' under the `.zattrs` entry. The nodes are + empty when anything is wrong. + """ + if not is_object(data): + return {}, {}, not_an_object(data) + doc = data + problems: list[ValidationProblem] = [ + ValidationProblem((key,), "missing required key", "missing_key") + for key in ("zarr_consolidated_format", "metadata") + if key not in doc + ] + problems.extend(unexpected_keys(frozenset({"zarr_consolidated_format", "metadata"}), doc)) + problems.extend(check_literal(doc, "zarr_consolidated_format", 1)) + refined: dict[str, JSONValue] = {} + if "metadata" in doc: + entries = doc["metadata"] + if not is_object(entries) or not all(isinstance(k, str) for k in entries): + problems.append( + ValidationProblem( + ("metadata",), "expected an object with string keys", "invalid_type" + ) + ) + else: + for key, value in entries.items(): + if not isinstance(key, str): + continue + refine = ( + refine_user_data + if key.rsplit("/", 1)[-1] == ZARR_V2_ATTRIBUTES_STORE_KEY + else refine_json + ) + entry, found = refine(value, ("metadata", key)) + problems.extend(found) + refined[key] = entry + nodes: dict[str, ZarrV2NodeMetadata] = {} + by_path: dict[str, dict[str, str]] = {} + if len(problems) == 0: + # A second key for one file is one problem among the nodes': every + # node is still read, so a user sees everything at once. + by_path, doubled = _entries_by_path(refined) + problems.extend(doubled) + for path, names in by_path.items(): + node, found = _read_node_v2(path, names, refined, context) + problems.extend(found) + if node is not None: + nodes[path] = node + problems.extend(_hierarchy_problems_v2(by_path)) + if len(problems) != 0: + nodes = {} + return refined, nodes, with_input(problems, doc) + + +def _hierarchy_problems_v2(by_path: Mapping[str, Mapping[str, str]]) -> list[ValidationProblem]: + """What keeps the nodes a `.zmetadata` holds from making a hierarchy, as `hierarchy_problems` judges one: a node below an array at its entry, a group missing above a node at the entry it would have; an orphan `.zattrs` is no node.""" + # The root is a group unless an entry says otherwise: a `.zmetadata` + # need not list its own `.zgroup`. + nodes: dict[str, NodeType] = {"/": "group"} + for path, names in by_path.items(): + if ZARR_V2_ARRAY_METADATA_STORE_KEY in names: + nodes[f"/{path}"] = "array" + elif ZARR_V2_GROUP_METADATA_STORE_KEY in names: + nodes[f"/{path}"] = "group" + problems: list[ValidationProblem] = [] + for found in hierarchy_problems(nodes): + at = found.loc[0] + path = at[1:] if isinstance(at, str) else "" + names = by_path.get(path, {}) + key = names.get( + ZARR_V2_ARRAY_METADATA_STORE_KEY, + names.get( + ZARR_V2_GROUP_METADATA_STORE_KEY, + f"{path}/{ZARR_V2_GROUP_METADATA_STORE_KEY}" + if path != "" + else ZARR_V2_GROUP_METADATA_STORE_KEY, + ), + ) + problems.append(ValidationProblem(("metadata", key), found.message, found.kind)) + return problems + + +def _read_node_v2( + path: str, names: Mapping[str, str], entries: Mapping[str, JSONValue], context: Context +) -> tuple[ZarrV2NodeMetadata | None, list[ValidationProblem]]: + """The node at `path`, read from its `.zarray` or `.zgroup` entry merged with its `.zattrs`, and every problem, located under the entries; None and no problem when there is only a `.zattrs`.""" + zarray = names.get(ZARR_V2_ARRAY_METADATA_STORE_KEY) + zgroup = names.get(ZARR_V2_GROUP_METADATA_STORE_KEY) + zattrs = names.get(ZARR_V2_ATTRIBUTES_STORE_KEY) + if zarray is not None and zgroup is not None: + return None, [ + ValidationProblem( + ("metadata", zgroup), + f"a node is an array or a group, not both: a {ZARR_V2_ARRAY_METADATA_STORE_KEY} is at " + f"{path!r} too", + "invalid_value", + ) + ] + key = zarray if zarray is not None else zgroup + if key is None: + return None, [] + document = entries[key] + if not is_json_object(document): + return None, [ + ValidationProblem( + ("metadata", key), f"expected an object, got {shown(document)}", "invalid_type" + ) + ] + merged = dict(document) + if "attributes" in merged: + return None, [ + ValidationProblem( + ("metadata", key, "attributes"), "unexpected document member", "invalid_value" + ) + ] + if zattrs is not None: + merged["attributes"] = entries[zattrs] + + def located(found: tuple[ValidationProblem, ...]) -> list[ValidationProblem]: + placed: list[ValidationProblem] = [] + for problem in found: + if zattrs is not None and problem.loc[:1] == ("attributes",): + placed.append( + dataclasses.replace(problem, loc=("metadata", zattrs, *problem.loc[1:])) ) else: - # Each entry is the document its key names: a `.zattrs` is - # user data, and any other is JSON by RFC 8259. - for key, value in cast("Mapping[str, object]", entries).items(): - refine = ( - refine_user_data - if key.rsplit("/", 1)[-1] == ZARR_V2_ATTRIBUTES_STORE_KEY - else refine_json - ) - entry, found = refine(value, ("metadata", key)) - problems.extend(found) - refined[key] = entry - if len(problems) != 0: - raise MetadataValidationError(problems) - return cls(metadata=refined) + placed.append(dataclasses.replace(problem, loc=("metadata", key, *problem.loc))) + return placed - @classmethod - def from_key_value(cls, mapping: Mapping[StoreKey, bytes]) -> ZarrV2ConsolidatedMetadata: - return cls.from_json(load_store_json(mapping, ZARR_V2_CONSOLIDATED_METADATA_STORE_KEY)) + if zarray is not None: + reading, members = read_array_v2(merged, context) + if members is None: + return None, located(reading.problems) + held = refined_object(merged) + if "dimension_separator" not in held: + held["dimension_separator"] = "." + return ZarrV2ArrayMetadata._of(held, context, reading, members), [] # pyright: ignore[reportPrivateUsage] + found = validate_group_metadata_v2(merged, context=context) + if len(found) != 0: + return None, located(found) + document = refined_object(merged) + attributes: dict[str, JSONValue] | UNSET = ( + object_at(document, "attributes") if "attributes" in document else UNSET + ) + return ZarrV2GroupMetadata._of(document, context, attributes), [] # pyright: ignore[reportPrivateUsage] - def to_key_value( - self, *, indent: int | str | None = None - ) -> Mapping[ZarrV2ConsolidatedMetadataStoreKey, bytes]: - # A model built by hand is not validated: it is written only as - # `from_json` reads it, and every problem is raised. - document = ZarrV2ConsolidatedMetadata.from_json(self.to_json()).to_json() - return {ZARR_V2_CONSOLIDATED_METADATA_STORE_KEY: dump_store_json(document, indent=indent)} + +def _v2_attributes(document: ZarrV2GroupMetadataJSON) -> dict[str, JSONValue] | UNSET: + """The attributes of the v2 group model of `document`, which `parse_group_metadata_v2` gave: `UNSET` when it holds none.""" + return dict(document["attributes"]) if "attributes" in document else UNSET diff --git a/packages/zarr-metadata/src/zarr_metadata/model/_json_schema.py b/packages/zarr-metadata/src/zarr_metadata/model/_json_schema.py new file mode 100644 index 0000000000..fc7c9e8358 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/model/_json_schema.py @@ -0,0 +1,101 @@ +"""The JSON Schema of a v3 `zarr.json`, as a scope reads one.""" + +from __future__ import annotations + +from typing import TYPE_CHECKING, cast + +from zarr_metadata._typed_json import Schemas +from zarr_metadata.v3._definition import DataTypeDefinition, field_schemas, written_name +from zarr_metadata.v3._registry import CORE_AND_EXTENSIONS, Context, ZarrV3Context, scoped +from zarr_metadata.v3.array import ZarrV3ArrayMetadataJSON +from zarr_metadata.v3.consolidated import ( + ZARR_V3_CONSOLIDATED_METADATA_KEY, + ZarrV3ConsolidatedMetadataJSON, +) +from zarr_metadata.v3.group import ZarrV3GroupMetadataJSON + +if TYPE_CHECKING: + from zarr_metadata._common import JSONValue + from zarr_metadata._typed_json import JSONSchema, SchemaLeaf + + +def node_metadata_json_schema_v3(*, context: ZarrV3Context | None = None) -> JSONSchema: + """The JSON Schema of a v3 `zarr.json` read in `context`: an array document or a group document, as `validate_node_metadata_v3` reads one, but for the rules. + + For an editor that validates a `zarr.json` as it is written, or a + validator in another language. JSON Schema draft 2020-12, as + `json_schema` writes one. Each extension point is a field as a scope + reads one in `context`: one a definition in scope reads, with the + configuration its definition declares, or a name none of them claims. The fill value is the JSON shape + the data type's definition declares for one -- an `int8`'s an integer + in [-128, 127] -- when the document names a data type in scope. A + group's `consolidated_metadata` holds array and group documents, by + path; a `null` one, which zarr-python 3.0 and 3.1 wrote, is refused, as + the validator refuses it. Each document is in `$defs` under the name of its + TypedDict: `ZarrV3ArrayMetadataJSON` is an array's alone. + + A JSON Schema says what each member is, and what the rules say of + members read together is not in it: one dimension name per dimension + of the shape, a chunk grid that fits the shape, codecs in the order a + pipeline takes them, each against the chunk it is handed, the + hierarchy the documents of consolidated metadata make below their + group, and what a definition's `rules` say. So a document it accepts may still have a + problem, and a JSON document `validate_node_metadata_v3` finds none + with, it accepts. A validator reads JSON as a parser gives it, arrays + as lists: a model's `to_json` writes tuples, which a Python validator + does not take for arrays. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + schemas = Schemas(_documents(scope)) + array = schemas.of(ZarrV3ArrayMetadataJSON) + group = schemas.of(ZarrV3GroupMetadataJSON) + return schemas.document({"anyOf": [array, group]}) + + +def _documents(context: Context) -> SchemaLeaf: + """The schema leaf of the documents read in `context`: an array's fill value held to its data type, and a group's consolidated metadata; each field as `context` reads one.""" + fields = field_schemas(context) + + def leaf(annotation: object, schemas: Schemas) -> JSONSchema | None: + if annotation is ZarrV3ArrayMetadataJSON: + return schemas.defined( + annotation, annotation.__name__, lambda: _array(context, schemas) + ) + if annotation is ZarrV3GroupMetadataJSON: + return schemas.defined(annotation, annotation.__name__, lambda: _group(schemas)) + return fields(annotation, schemas) + + return leaf + + +def _array(context: Context, schemas: Schemas) -> JSONSchema: + """An array document: its TypedDict, and its fill value held to the data type it names, for each data type in scope.""" + schema = schemas.object_of(ZarrV3ArrayMetadataJSON) + held: list[JSONValue] = [] + for definition in context.definitions(): + if not isinstance(definition, DataTypeDefinition): + continue + fill_value = schemas.of(definition.fill_value) + if len(fill_value) == 0: + continue # a data type that says nothing of its fill value takes any JSON + name = written_name(definition) + named: JSONSchema = { + "anyOf": [name, {"type": "object", "properties": {"name": name}, "required": ["name"]}] + } + condition: JSONSchema = { + "if": {"properties": {"data_type": named}, "required": ["data_type"]}, + "then": {"properties": {"fill_value": fill_value}}, + } + held.append(condition) + return schema if len(held) == 0 else {**schema, "allOf": held} + + +def _group(schemas: Schemas) -> JSONSchema: + """A group document: its TypedDict, and the consolidated metadata the model reads.""" + schema = schemas.object_of(ZarrV3GroupMetadataJSON) + properties = cast("dict[str, JSONValue]", schema.get("properties", {})) + consolidated: JSONSchema = schemas.of(ZarrV3ConsolidatedMetadataJSON) + return {**schema, "properties": {**properties, ZARR_V3_CONSOLIDATED_METADATA_KEY: consolidated}} + + +__all__ = ["node_metadata_json_schema_v3"] diff --git a/packages/zarr-metadata/src/zarr_metadata/model/_keyed.py b/packages/zarr-metadata/src/zarr_metadata/model/_keyed.py new file mode 100644 index 0000000000..f131c7748f --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/model/_keyed.py @@ -0,0 +1,22 @@ +"""What every model compares and hashes by: a key its constructor computes once of everything the model shows.""" + + +class Keyed: + """A model whose `==` and `hash` compare a key computed once, when it is built, of what its document means. + + Each model computes its own key in `_key_of` and stores it; two models + are equal when they are of one class and their keys are, so a model of + another class, or anything else, is never equal to one. + """ + + __slots__ = ("_key",) + + _key: tuple[object, ...] + + def __eq__(self, other: object) -> bool: + if type(other) is not type(self): + return NotImplemented + return self._key == other._key + + def __hash__(self) -> int: + return hash(self._key) diff --git a/packages/zarr-metadata/src/zarr_metadata/model/_repair.py b/packages/zarr-metadata/src/zarr_metadata/model/_repair.py new file mode 100644 index 0000000000..f0e8b91062 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/model/_repair.py @@ -0,0 +1,346 @@ +"""Repairing v3 metadata that a known writer bug made invalid. + +The readers are strict: a document the spec does not allow has problems, +whoever wrote it. Some writers have written documents the spec does not +allow, in ways that say plainly what they meant. Repairing one is a step +of its own, before a strict read, which a caller asks for by name: +`read_repaired_node_metadata_v3` rather than `read_node_metadata_v3`. + +What is repaired is a closed set, each member of it a writer's bug: the +shape of the JSON it wrote is a TypedDict, which a document must have for +the repair to apply, and the repair makes it correct JSON. Anything else +-- a missing `data_type`, a chunk length of 0 along a dimension that is +not empty -- is left as it is, for the strict read to report. +""" + +from __future__ import annotations + +from collections.abc import Mapping +from dataclasses import dataclass +from typing import Annotated, Literal, TypeAlias, TypedDict, cast + +from annotated_types import Ge + +from zarr_metadata._common import ( + JSONValue, +) +from zarr_metadata._json import ( + JSON_DEPTH, + MetadataValidationError, + ValidationProblem, + is_json_object, + is_object, +) +from zarr_metadata._typed_json import Loc, check +from zarr_metadata.model._group import ( + ZarrV2ConsolidatedMetadata, + ZarrV3NodeMetadataReading, + read_node_metadata_v3, +) +from zarr_metadata.v2.definition import CORE_V2 +from zarr_metadata.v2.group import ZARR_V2_GROUP_METADATA_STORE_KEY +from zarr_metadata.v3._registry import CORE_AND_EXTENSIONS, ZarrV2Context, ZarrV3Context, scoped +from zarr_metadata.v3.consolidated import ZARR_V3_CONSOLIDATED_METADATA_KEY + +RepairKind: TypeAlias = Literal[ + "zero_chunk_length", "null_consolidated_metadata", "consolidated_metadata_in_zgroup_entry" +] +"""Each writer bug a repair undoes, by name.""" + + +@dataclass(frozen=True, slots=True) +class Repair: + """One change a repair made to a document: where, which bug it undid, and what it did.""" + + loc: Loc + """Where the change is, in the document: the value changed, or the key removed.""" + kind: RepairKind + """The writer bug the change undoes.""" + message: str + """What was written, by which writer, and what it became.""" + + +_Length = Annotated[int, Ge(0)] + + +class ZarrV3ZeroChunkRegularGridConfigurationJSON(TypedDict): + """A `regular` grid's configuration as zarr-python 3.0 and 3.1 wrote it: chunk lengths that may be 0.""" + + chunk_shape: tuple[_Length, ...] + + +class ZarrV3ZeroChunkRegularGridJSON(TypedDict): + """A `regular` chunk grid whose chunk lengths may be 0.""" + + name: Literal["regular"] + configuration: ZarrV3ZeroChunkRegularGridConfigurationJSON + + +class ZarrV3ZeroChunkArrayMetadataJSON(TypedDict): + """The members of an array document the `zero_chunk_length` repair reads. + + zarr-python 3.0 and 3.1 wrote a chunk length of 0 along a dimension of + length 0, which the regular grid refuses: "Chunk sizes must be greater + than zero". A chunk along an empty dimension holds nothing whatever its + length, so the repair writes 1 there, as zarr-python has since + (https://github.com/zarr-developers/zarr-python/pull/4328). + """ + + node_type: Literal["array"] + shape: tuple[_Length, ...] + chunk_grid: ZarrV3ZeroChunkRegularGridJSON + + +class ZarrV3NullConsolidatedGroupMetadataJSON(TypedDict): + """The members of a group document the `null_consolidated_metadata` repair reads. + + zarr-python 3.0 and 3.1 wrote `"consolidated_metadata": null` on a group + they had not consolidated, which the convention does not allow: the member is + an object, or absent. The repair removes it, which is what the writer + meant. + """ + + node_type: Literal["group"] + consolidated_metadata: None + + +def repair_node_metadata_v3(value: object) -> tuple[object, tuple[Repair, ...]]: + """`value`, a v3 `zarr.json`, with each known writer bug in it undone, and what was changed. + + Each document consolidated metadata holds is repaired too. What no + repair applies to is left as it is, and `value` is not changed: a + repaired document is a new one, sharing what it did not change with + `value`. Repairing a document with none of the bugs gives it back, + and no repairs. + """ + return _repaired(value, ()) + + +def _repaired(value: object, at: Loc) -> tuple[object, tuple[Repair, ...]]: + if not is_json_object(value) or len(at) >= JSON_DEPTH: + # Not a document, or past the levels a reader walks, which the + # strict read reports. + return value, () + original = value + repairs: list[Repair] = [] + document = _zero_chunk_length(original, at, repairs) + document = _null_consolidated_metadata(document, at, repairs) + document = _consolidated(document, at, repairs) + return (original if len(repairs) == 0 else dict(document)), tuple(repairs) + + +def _members(document: Mapping[str, object], keys: tuple[str, ...]) -> dict[str, object]: + """The members of `document` a repair reads: only those, so the rest is not walked.""" + return {key: document[key] for key in keys if key in document} + + +def _zero_chunk_length( + document: Mapping[str, object], at: Loc, repairs: list[Repair] +) -> Mapping[str, object]: + array, problems = check( + _members(document, ("node_type", "shape", "chunk_grid")), ZarrV3ZeroChunkArrayMetadataJSON + ) + if array is None or len(problems) != 0: + return document + shape = array["shape"] + chunk_shape = array["chunk_grid"]["configuration"]["chunk_shape"] + zeros = [axis for axis, length in enumerate(chunk_shape) if length == 0] + if len(zeros) == 0 or len(chunk_shape) != len(shape) or any(shape[a] != 0 for a in zeros): + # Nothing to repair, or a 0 that is no writer's bug: the strict + # read reports it. + return document + repairs.extend( + Repair( + (*at, "chunk_grid", "configuration", "chunk_shape", axis), + "zero_chunk_length", + "a chunk length of 0 along a dimension of length 0, as zarr-python 3.0 and 3.1 " + "wrote it, written as 1", + ) + for axis in zeros + ) + grid = cast("Mapping[str, object]", document["chunk_grid"]) + configuration = cast("Mapping[str, object]", grid["configuration"]) + repaired = tuple(max(length, 1) for length in chunk_shape) + return { + **document, + "chunk_grid": {**grid, "configuration": {**configuration, "chunk_shape": repaired}}, + } + + +def _null_consolidated_metadata( + document: Mapping[str, object], at: Loc, repairs: list[Repair] +) -> Mapping[str, object]: + group, problems = check( + _members(document, ("node_type", ZARR_V3_CONSOLIDATED_METADATA_KEY)), + ZarrV3NullConsolidatedGroupMetadataJSON, + ) + if group is None or len(problems) != 0: + return document + repairs.append( + Repair( + (*at, ZARR_V3_CONSOLIDATED_METADATA_KEY), + "null_consolidated_metadata", + "a consolidated_metadata of null, as zarr-python 3.0 and 3.1 wrote it, removed", + ) + ) + return {key: item for key, item in document.items() if key != ZARR_V3_CONSOLIDATED_METADATA_KEY} + + +def _consolidated( + document: Mapping[str, object], at: Loc, repairs: list[Repair] +) -> Mapping[str, object]: + """`document` with each document its consolidated metadata holds repaired, where it holds any.""" + member = document.get(ZARR_V3_CONSOLIDATED_METADATA_KEY) + if not is_object(member): + return document + envelope = member + entries = envelope.get("metadata") + if not is_object(entries): + return document + held: dict[object, object] = {} + found = len(repairs) + for key, entry in entries.items(): + if isinstance(key, str): + entry, inside = _repaired( + entry, (*at, ZARR_V3_CONSOLIDATED_METADATA_KEY, "metadata", key) + ) + repairs.extend(inside) + held[key] = entry + if len(repairs) == found: + return document + return {**document, ZARR_V3_CONSOLIDATED_METADATA_KEY: {**envelope, "metadata": held}} + + +@dataclass(frozen=True, slots=True) +class ZarrV3RepairedNodeMetadataReading: + """A v3 `zarr.json` read after its known writer bugs were undone: the strict reading of the repaired document, and the repairs.""" + + reading: ZarrV3NodeMetadataReading + """The repaired document, as `read_node_metadata_v3` reads it: its problems are the repaired document's, and its model when there are none.""" + repairs: tuple[Repair, ...] + """What was changed to make the document that was read.""" + + +def read_repaired_node_metadata_v3( + value: object, *, context: ZarrV3Context | None = None +) -> ZarrV3RepairedNodeMetadataReading: + """`value`, a v3 `zarr.json`, read in `context` as `read_node_metadata_v3` reads it, once `repair_node_metadata_v3` has undone each known writer bug in it. + + For a reader of stores other writers made, which asks for repairs by + calling this rather than `read_node_metadata_v3`. Whatever no repair + applies to is read as it is, and reported as `read_node_metadata_v3` + reports it. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + repaired, repairs = repair_node_metadata_v3(value) + return ZarrV3RepairedNodeMetadataReading( + read_node_metadata_v3(repaired, context=scope), repairs + ) + + +class ZarrV2ZGroupWithConsolidatedMetadataJSON(TypedDict): + """A `.zgroup` entry of a `.zmetadata` as zarr-python 3.x writes one below the root: with a `consolidated_metadata` member, which a v2 group document does not take.""" + + zarr_format: Literal[2] + consolidated_metadata: Mapping[str, JSONValue] + + +def repair_consolidated_metadata_v2(value: object) -> tuple[object, tuple[Repair, ...]]: + """`value`, a v2 `.zmetadata`, with each known writer bug in it undone, and what was changed. + + zarr-python 3.x writes a `consolidated_metadata` member into each + `.zgroup` entry below the root, which is removed; the root's is left, + since no writer puts one there. What no repair + applies to is left as it is, and `value` is not changed; a document + with none of the bugs is given back, and no repairs. + """ + if not is_json_object(value): + return value, () + document = value + entries = document.get("metadata") + if not is_object(entries): + return value, () + repairs: list[Repair] = [] + held: dict[object, object] = {} + for key, entry in entries.items(): + if isinstance(key, str): + path, _, name = key.rpartition("/") + # Below the root only: no writer puts the member in the root's .zgroup. + if name == ZARR_V2_GROUP_METADATA_STORE_KEY and path.strip("/") != "": + entry = _without_consolidated_metadata(entry, ("metadata", key), repairs) + held[key] = entry + if len(repairs) == 0: + return value, () + return {**document, "metadata": held}, tuple(repairs) + + +def _without_consolidated_metadata(entry: object, at: Loc, repairs: list[Repair]) -> object: + """`entry`, a `.zgroup` entry, without the member zarr-python 3.x writes into it, when it is one such.""" + if not is_json_object(entry): + return entry + group = entry + shaped, problems = check( + _members(group, ("zarr_format", ZARR_V3_CONSOLIDATED_METADATA_KEY)), + ZarrV2ZGroupWithConsolidatedMetadataJSON, + ) + if shaped is None or len(problems) != 0: + return entry + repairs.append( + Repair( + (*at, ZARR_V3_CONSOLIDATED_METADATA_KEY), + "consolidated_metadata_in_zgroup_entry", + "a consolidated_metadata, as zarr-python 3.x writes into a .zgroup entry of a " + ".zmetadata, removed", + ) + ) + kept: dict[str, object] = { + key: item for key, item in group.items() if key != ZARR_V3_CONSOLIDATED_METADATA_KEY + } + return kept + + +@dataclass(frozen=True, slots=True) +class ZarrV2RepairedConsolidatedMetadataReading: + """A v2 `.zmetadata` read after its known writer bugs were undone: the repaired document's problems, its model when there are none, and the repairs.""" + + problems: tuple[ValidationProblem, ...] + """Every problem of the repaired document.""" + metadata: ZarrV2ConsolidatedMetadata | None + """The repaired document's model, when it has no problem; None otherwise.""" + repairs: tuple[Repair, ...] + """What was changed to make the document that was read.""" + + +def read_repaired_consolidated_metadata_v2( + value: object, *, context: ZarrV2Context | None = None +) -> ZarrV2RepairedConsolidatedMetadataReading: + """`value`, a v2 `.zmetadata`, read in `context` as `ZarrV2ConsolidatedMetadata` reads it, once `repair_consolidated_metadata_v2` has undone each known writer bug in it. + + For a reader of stores other writers made, which asks for repairs by + calling this rather than the strict model. Whatever no repair applies + to is read as it is, and reported as the strict read reports it. + """ + scope = scoped(context, CORE_V2) + repaired, repairs = repair_consolidated_metadata_v2(value) + try: + model: ZarrV2ConsolidatedMetadata | None = ZarrV2ConsolidatedMetadata(repaired, scope) + except MetadataValidationError as error: + return ZarrV2RepairedConsolidatedMetadataReading(error.problems, None, repairs) + return ZarrV2RepairedConsolidatedMetadataReading((), model, repairs) + + +__all__ = [ + "Repair", + "RepairKind", + "ZarrV2RepairedConsolidatedMetadataReading", + "ZarrV2ZGroupWithConsolidatedMetadataJSON", + "ZarrV3NullConsolidatedGroupMetadataJSON", + "ZarrV3RepairedNodeMetadataReading", + "ZarrV3ZeroChunkArrayMetadataJSON", + "ZarrV3ZeroChunkRegularGridConfigurationJSON", + "ZarrV3ZeroChunkRegularGridJSON", + "read_repaired_consolidated_metadata_v2", + "read_repaired_node_metadata_v3", + "repair_consolidated_metadata_v2", + "repair_node_metadata_v3", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/model/_validation.py b/packages/zarr-metadata/src/zarr_metadata/model/_validation.py index d1ae10b83a..bf7c68d38c 100644 --- a/packages/zarr-metadata/src/zarr_metadata/model/_validation.py +++ b/packages/zarr-metadata/src/zarr_metadata/model/_validation.py @@ -7,13 +7,14 @@ name nothing in the scope claims is left unjudged. A v3 fill value is judged against the data type it names, the chunk grid against the shape, and the codecs as a pipeline, each against the chunk it is -handed. Each concept -gets a `validate_*` function returning every problem found, an `is_*` -type guard, and a `parse_*` function that narrows or raises -`MetadataValidationError`. The guards are `TypeGuard`s, -not `TypeIs`: True narrows a value to its document type, and False says -nothing about its type, since a value can be well typed and still not a -valid document. +handed. Each concept gets a `validate_*` function returning every +problem found, an `is_*` type guard, and a `parse_*` function that +narrows or raises `MetadataValidationError`; a v3 array document also +gets `read_array_metadata_v3`, one read that returns what it read, the +problems, and the model when there are none. +The guards are `TypeGuard`s, not `TypeIs`: True narrows a value to its +document type, and False says nothing about its type, since a value can +be well typed and still not a valid document. Every `ValidationProblem` carries a machine-readable `kind` alongside its human-readable `message`, so consumers can dispatch on the failure mode @@ -23,40 +24,83 @@ from __future__ import annotations +import dataclasses import json from collections.abc import Mapping, Sequence -from typing import Any, Final, TypeGuard, TypeVar, cast +from dataclasses import dataclass +from typing import ( + TYPE_CHECKING, + Any, + Final, + Protocol, + TypeGuard, + TypeVar, + cast, + get_args, + get_origin, +) from zarr_metadata._json import ( MetadataValidationError, ValidationProblem, arrays_to_tuples, + is_object, + is_tuple, + nested_past_the_levels, + not_an_object, + outside_of, + refine_json, refine_user_data, - validate_json, + shown_key, + with_input, + within, ) from zarr_metadata._json import is_canonical_json as _is_canonical_json -from zarr_metadata._json import prefixed as _prefix +from zarr_metadata._sentinel import UNSET +from zarr_metadata._typed_json import typeddict_keys from zarr_metadata.v2.array import ZarrV2ArrayMetadataJSON +from zarr_metadata.v2.definition import ( + CORE_V2, + ZarrV2CodecDefinition, + ZarrV2DataTypeDefinition, + resolve_codec_v2, + resolve_dtype_v2, +) from zarr_metadata.v2.group import ZarrV2GroupMetadataJSON from zarr_metadata.v3._definition import ( Chunk, ChunkGridDefinition, ChunkKeyEncodingDefinition, - CodecDefinition, DataTypeDefinition, Definition, Lengths, - Resolved, + ResolvedField, StorageTransformerDefinition, chunk_grid_lengths, + field_kind, + fields_of, fill_value_problems, resolve, ) -from zarr_metadata.v3._pipeline import read_pipeline -from zarr_metadata.v3._registry import CORE_AND_EXTENSIONS, Context +from zarr_metadata.v3._pipeline import Stage, read_pipeline +from zarr_metadata.v3._registry import ( + CORE_AND_EXTENSIONS, + Context, + ZarrV2Context, + ZarrV3Context, + scoped, +) from zarr_metadata.v3.array import ZarrV3ArrayMetadataJSON from zarr_metadata.v3.group import ZarrV3GroupMetadataJSON +if TYPE_CHECKING: + from collections.abc import Iterator + + from zarr_metadata._common import JSONValue + from zarr_metadata._typed_json import Loc + from zarr_metadata.model._array import ZarrV2ArrayMetadata, ZarrV3ArrayMetadata + from zarr_metadata.v2.array import ZarrV2ArrayDimensionSeparator, ZarrV2ArrayOrder + # The standard top-level keys of a v3 array metadata document. Anything outside # this set is an extension field. Built from the TypedDict's required/optional # key sets (which resolve inherited keys, unlike `__annotations__`). @@ -103,7 +147,7 @@ ) -def _missing_keys( +def missing_keys( required: frozenset[str], doc: Mapping[object, object] ) -> tuple[ValidationProblem, ...]: """One `missing_key` problem per required key absent from `doc`.""" @@ -113,35 +157,34 @@ def _missing_keys( ) -def _unexpected_keys( +def unexpected_keys( allowed: frozenset[str], doc: Mapping[object, object] ) -> tuple[ValidationProblem, ...]: - """One problem per member outside a closed document's declared shape.""" + """One problem per member outside a closed document's declared shape: `unknown_key`, as a closed TypedDict's checker reports one, so a caller who tolerates a member another writer added can tell it from a wrong value.""" problems: list[ValidationProblem] = [] for key in doc: if not isinstance(key, str): problems.append( - ValidationProblem((), f"non-string document key {key!r}", "invalid_type") + ValidationProblem((), f"non-string document key {shown_key(key)}", "invalid_type") ) elif key not in allowed: - problems.append( - ValidationProblem((key,), "unexpected document member", "invalid_value") - ) + problems.append(ValidationProblem((key,), f"unexpected key {key!r}", "unknown_key")) return tuple(problems) -def _check_literal( +def check_literal( doc: Mapping[object, object], key: str, expected: object ) -> tuple[ValidationProblem, ...]: - """One `invalid_value` problem if `doc[key]` is present but not `expected`.""" - if key in doc and (type(doc[key]) is not type(expected) or doc[key] != expected): - return ( - ValidationProblem((key,), f"expected {expected!r}, got {doc[key]!r}", "invalid_value"), - ) + """One problem if `doc[key]` is present but not `expected`: of its type, when it is not of `expected`'s JSON type, else of its value; a value is judged as refined, so a `StrEnum` member is its string.""" + if key not in doc: + return () + value, found = refine_json(doc[key], (key,)) + if len(found) != 0 or type(value) is not type(expected) or value != expected: + return (outside_of((key,), doc[key], (expected,)),) return () -def _validate_other_members( +def other_members_problems( doc: Mapping[object, object], standard_keys: frozenset[str], *, @@ -152,18 +195,33 @@ def _validate_other_members( For a document open to other members: v3 extension fields, and the members a v2 array's readers ignore. """ + return other_members(doc, standard_keys, additional_reserved_keys=additional_reserved_keys)[1] + + +def other_members( + doc: Mapping[object, object], + standard_keys: frozenset[str], + *, + additional_reserved_keys: frozenset[str] = frozenset(), + at: Loc = (), +) -> tuple[dict[str, JSONValue], tuple[ValidationProblem, ...]]: + """Each member outside `standard_keys` refined to JSON, and every problem, as `other_members_problems` finds them; a member that is not JSON is left out.""" + members: dict[str, JSONValue] = {} problems: list[ValidationProblem] = [] reserved_keys = standard_keys | additional_reserved_keys for key, value in doc.items(): if not isinstance(key, str): problems.append( - ValidationProblem((), f"non-string top-level key {key!r}", "invalid_type") + ValidationProblem((), f"non-string top-level key {shown_key(key)}", "invalid_type") ) continue if key in reserved_keys: continue - problems.extend(_prefix(key, validate_json(value))) - return tuple(problems) + refined, found = refine_json(value, (*at, key)) + problems.extend(within(found, at)) + if len(found) == 0: + members[key] = refined + return members, tuple(problems) def _is_array(value: object) -> TypeGuard[Sequence[object]]: @@ -188,8 +246,8 @@ def _is_int_sequence(value: object) -> TypeGuard[Sequence[int]]: ) -def _dimension_lengths( - doc: Mapping[object, object], key: str +def dimension_lengths( + doc: Mapping[object, object] | Mapping[str, object], key: str ) -> tuple[tuple[int, ...] | None, tuple[ValidationProblem, ...]]: """The dimension lengths `doc` holds at `key` (`shape`, `chunks`), and every problem with them. @@ -198,48 +256,27 @@ def _dimension_lengths( """ if key not in doc: return None, () - value = doc[key] + value, found = refine_json(doc[key], (key,)) + if len(found) != 0: + # An integer JSON text does not hold, or a value nested too deep. + return None, found if not _is_int_sequence(value): - return None, (ValidationProblem((key,), "expected a sequence of int", "invalid_type"),) + return None, (ValidationProblem((key,), "expected an array of integers", "invalid_type"),) if any(item < 0 for item in value): return None, (ValidationProblem((key,), "expected non-negative integers", "invalid_value"),) return tuple(value), () -def _is_dtype_v2(value: object) -> bool: - """Whether `value` is shaped like a v2 dtype: a string or field records. - - A field record is a `(name, dtype)` or `(name, dtype, shape)` sequence, - where `dtype` is itself a string or nested field records and `shape` is a - sequence of int. The string content is NOT interpreted — whether the - string names a real dtype is domain validity, not structure. - """ - if isinstance(value, str): - return True - if not _is_array(value): - return False - for record in value: - if not _is_array(record) or len(record) not in (2, 3): - return False - if not isinstance(record[0], str): - return False - if not _is_dtype_v2(record[1]): - return False - if len(record) == 3 and not _is_int_sequence(record[2]): - return False - return True - - def _is_canonical_dtype_v2(value: object) -> bool: """Whether a validated v2 dtype uses the tuple-backed public representation.""" if isinstance(value, str): return True - if not isinstance(value, tuple): + if not is_tuple(value): return False - for record in cast("tuple[object, ...]", value): - if not isinstance(record, tuple): + for record in value: + if not is_tuple(record): return False - fields = cast("tuple[object, ...]", record) + fields = record if not _is_canonical_dtype_v2(fields[1]): return False if len(fields) == 3 and not isinstance(fields[2], tuple): @@ -252,34 +289,29 @@ def _is_canonical_metadata_field_v3(value: object) -> bool: return isinstance(value, (str, dict)) -def _is_canonical_array_metadata_v3(value: object) -> bool: +def is_canonical_array_metadata_v3(value: object) -> bool: """Whether a validated v3 array document matches `ZarrV3ArrayMetadataJSON` at runtime.""" - if not isinstance(value, dict): + if not is_object(value) or not isinstance(value, dict): return False - doc = cast("dict[str, object]", value) - if not isinstance(doc["shape"], tuple) or not isinstance(doc["codecs"], tuple): + doc = value + codecs = doc["codecs"] + if not isinstance(doc["shape"], tuple) or not is_tuple(codecs): return False - if "storage_transformers" in doc and not isinstance(doc["storage_transformers"], tuple): + transformers = doc.get("storage_transformers", ()) + if not is_tuple(transformers): return False if "dimension_names" in doc and not isinstance(doc["dimension_names"], tuple): return False if not all(_is_canonical_metadata_field_v3(doc[key]) for key, _ in _EXTENSION_POINTS_V3): return False - if not all( - _is_canonical_metadata_field_v3(item) for item in cast("tuple[object, ...]", doc["codecs"]) - ): - return False - return "storage_transformers" not in doc or all( - _is_canonical_metadata_field_v3(item) - for item in cast("tuple[object, ...]", doc["storage_transformers"]) - ) + return all(_is_canonical_metadata_field_v3(item) for item in (*codecs, *transformers)) def _is_canonical_array_metadata_v2(value: object) -> bool: """Whether a validated v2 array document matches `ZarrV2ArrayMetadataJSON` at runtime.""" - if not isinstance(value, dict): + if not is_object(value) or not isinstance(value, dict): return False - doc = cast("dict[str, object]", value) + doc = value if not isinstance(doc["shape"], tuple) or not isinstance(doc["chunks"], tuple): return False if not _is_canonical_dtype_v2(doc["dtype"]): @@ -289,30 +321,11 @@ def _is_canonical_array_metadata_v2(value: object) -> bool: return False filters = doc["filters"] return filters is None or ( - isinstance(filters, tuple) - and all(isinstance(item, dict) for item in cast("tuple[object, ...]", filters)) - ) - - -def _is_codec_v2(value: object) -> bool: - """Whether `value` is shaped like a v2 codec config: a mapping with a string `id`.""" - return isinstance(value, Mapping) and isinstance( - cast("Mapping[object, object]", value).get("id"), str + is_tuple(filters) and all(isinstance(item, dict) for item in filters) ) -def _validate_codec_v2(value: object) -> tuple[ValidationProblem, ...]: - """Validate a v2 codec's required shape and JSON-valued configuration.""" - if not _is_codec_v2(value): - return ( - ValidationProblem( - (), "expected a codec configuration with a string 'id'", "invalid_type" - ), - ) - return validate_json(value) - - -def _validate_attributes(value: object) -> tuple[ValidationProblem, ...]: +def validate_attributes(value: object) -> tuple[ValidationProblem, ...]: """Validate an `attributes` value: a mapping with string keys. Returns a problem at `("attributes",)` if it is not, else `[]`. Shared by the @@ -321,64 +334,220 @@ def _validate_attributes(value: object) -> tuple[ValidationProblem, ...]: already-parent-relative `("attributes",)` loc, since it is only ever called with a document's `attributes` value. """ - if not isinstance(value, Mapping) or not all( - isinstance(k, str) for k in cast("Mapping[object, object]", value) - ): - return ( + return attributes_of(value)[1] + + +def members_past_the_levels( + doc: Mapping[object, object], at: Loc +) -> dict[object, ValidationProblem]: + """Each member of `doc`, a document that sits at `at`, that is a container past the levels a reader walks, with the problem it is, located in the document. + + A reader reports them and walks no further into them, as `refine_json` + of the whole document would judge them: a document a level before the + cap holds its scalars, and nothing else. + """ + past: dict[object, ValidationProblem] = {} + for key, item in doc.items(): + if isinstance(key, str): + problem = nested_past_the_levels(item, (*at, key)) + if problem is not None: + (past[key],) = within((problem,), at) + return past + + +def attributes_of( + value: object, at: Loc = () +) -> tuple[dict[str, JSONValue] | None, tuple[ValidationProblem, ...]]: + """An `attributes` value refined as user data, and every problem `validate_attributes` finds; None when it is not an object with string keys, or holds a value that is not JSON. + + `at` is where the document holding it sits in the one handed in, so + the levels a reader walks are counted from that one's root; the + problems are located in the document holding it. + """ + if not is_object(value) or not all(isinstance(k, str) for k in value): + return None, ( ValidationProblem( - ("attributes",), "expected a mapping with string keys", "invalid_type" + ("attributes",), "expected an object with string keys", "invalid_type" ), ) + attributes: dict[str, JSONValue] = {} problems: list[ValidationProblem] = [] - for key, item in cast("Mapping[str, object]", value).items(): - problems.extend(refine_user_data(item, ("attributes", key))[1]) - return tuple(problems) + for key, item in value.items(): + if not isinstance(key, str): + continue + refined, found = refine_user_data(item, (*at, "attributes", key)) + problems.extend(found) + attributes[key] = refined + return (attributes if len(problems) == 0 else None), within(problems, at) + +_MEMBERS_V3: Final = typeddict_keys(ZarrV3ArrayMetadataJSON).members -_EXTENSION_POINTS_V3: Final[tuple[tuple[str, type[Definition[Any]]], ...]] = ( - ("data_type", DataTypeDefinition), - ("chunk_grid", ChunkGridDefinition), - ("chunk_key_encoding", ChunkKeyEncodingDefinition), +_EXTENSION_POINTS_V3: Final[tuple[tuple[str, type[Definition[Any]]], ...]] = tuple( + (key, kind) + for key, (annotation, _) in _MEMBERS_V3.items() + if (kind := field_kind(annotation)) is not None ) -"""A v3 array document's single extension points, and the kind each is read as.""" +"""A v3 array document's single extension points, and the kind each is read as: each member its TypedDict annotates with a field alias.""" -_EXTENSION_LISTS_V3: Final[tuple[tuple[str, type[Definition[Any]]], ...]] = ( - ("codecs", CodecDefinition), - ("storage_transformers", StorageTransformerDefinition), +_EXTENSION_LISTS_V3: Final[tuple[tuple[str, type[Definition[Any]]], ...]] = tuple( + (key, kind) + for key, (annotation, _) in _MEMBERS_V3.items() + if get_origin(annotation) is tuple and (kind := field_kind(get_args(annotation)[0])) is not None ) -"""Its lists of extension points, and the kind each entry is read as.""" +"""Its lists of extension points, and the kind each entry is read as: each member it annotates as a tuple of a field alias.""" -def validate_array_metadata_v3( - value: object, *, context: Context = CORE_AND_EXTENSIONS -) -> tuple[ValidationProblem, ...]: - """Return every reason `value` is not a valid v3 array document. +@dataclass(frozen=True, slots=True) +class ZarrV3ArrayMetadataReading: + """A v3 array document as a scope read it, whatever it holds: each extension point, its codecs as a pipeline, every problem, and the model when there is none. - Its structure, and each extension point read through the definition - that claims its name in `context`: a gzip `level` out of range, a key a - codec's configuration does not declare. The fill value is judged - against the data type as `context` read it -- an `int8` fill value of - 300 -- and the chunk grid against the shape: a regular grid with a - chunk length for each of two dimensions, over an array of three. The - codecs are read as a pipeline: in order, each judged against the chunk - it is handed -- a `transpose` whose `order` has another number of - axes, a shard its inner chunks do not divide -- and a shard's inner - and index codecs too. - A name nothing in `context` claims is left unjudged, with any fill - value of it, and a codec of that name leaves the codec after it - handed a chunk nothing is known of. Unknown top-level keys are - allowed (they map to `extra_fields`); a reader must understand each - one that does not say `must_understand: false`, which the model - reports as `must_understand_fields`. + A field the document does not hold is `UNSET`, and a list of them it does + not hold as a list is empty. + """ + + data_type: ResolvedField[DataTypeDefinition[Any]] | UNSET = UNSET + """The data type, as the scope read it.""" + chunk_grid: ResolvedField[ChunkGridDefinition[Any]] | UNSET = UNSET + """The chunk grid, as the scope read it.""" + chunk_key_encoding: ResolvedField[ChunkKeyEncodingDefinition[Any]] | UNSET = UNSET + """The chunk key encoding, as the scope read it.""" + chunk: Chunk = dataclasses.field(default_factory=Chunk) + """The chunks the codecs are handed: the lengths the grid's chunks take along each axis of the shape, of the data type.""" + pipeline: tuple[Stage, ...] = () + """The codecs, read as a pipeline: each as the scope read it, with the chunk it is handed.""" + storage_transformers: tuple[ResolvedField[StorageTransformerDefinition[Any]], ...] = () + """The storage transformers, each as the scope read it.""" + problems: tuple[ValidationProblem, ...] = () + """Every reason the document is not a valid one.""" + metadata: ZarrV3ArrayMetadata | None = None + """The document's model, holding these fields, when there is no problem; None otherwise.""" + + def __reduce__(self) -> str | tuple[object, ...]: + # A reading that holds its model pickles and copies as the model + # does -- the pair, read once on load -- and comes back as that + # model's own reading, so one model per document still. + if self.metadata is not None: + return (reading_of, (self.metadata,)) + return object.__reduce__(self) + + def fields(self) -> Iterator[tuple[Loc, ResolvedField[Any]]]: + """Each field the document holds, as the scope read it, with where it sits in the document. + + The extension points, then each codec and storage transformer at its + index, each followed by the fields it holds, as `fields_of` gives + them: a shard's codecs, a struct's field types. `with_problems` + gives each with its problems. + """ + own: tuple[tuple[str, ResolvedField[Any] | UNSET], ...] = ( + ("data_type", self.data_type), + ("chunk_grid", self.chunk_grid), + ("chunk_key_encoding", self.chunk_key_encoding), + ) + for key, field in own: + if field is not UNSET: + yield from fields_of(field, (key,)) + for index, stage in enumerate(self.pipeline): + yield from fields_of(stage.codec, ("codecs", index)) + for index, transformer in enumerate(self.storage_transformers): + yield from fields_of(transformer, ("storage_transformers", index)) + + +class _HoldsReading(Protocol): + """A model: what holds the reading it was built from.""" + + @property + def reading(self) -> object: ... + + +def reading_of(model: _HoldsReading) -> object: + """The reading `model`, a model, holds: what a pickled reading that held a model is built again as.""" + return model.reading + + +@dataclass(frozen=True, slots=True) +class ArrayMembersV3: + """The members of a v3 array document a read found nothing wrong with, other than its fields, refined as the model holds them.""" + + shape: tuple[int, ...] + fill_value: JSONValue + dimension_names: tuple[str | None, ...] | UNSET + attributes: dict[str, JSONValue] + extra_fields: dict[str, JSONValue] + + +@dataclass(frozen=True, slots=True) +class ArrayMembersV2: + """The members of a v2 array document a read found nothing wrong with, other than its fields, refined as the model holds them.""" + + shape: tuple[int, ...] + chunks: tuple[int, ...] + fill_value: JSONValue + order: ZarrV2ArrayOrder + dimension_separator: ZarrV2ArrayDimensionSeparator + attributes: dict[str, JSONValue] | UNSET + extra_fields: dict[str, JSONValue] + + +@dataclass(frozen=True, slots=True) +class ZarrV2ArrayMetadataReading: + """A v2 array document as a scope read it, whatever it holds: its dtype, compressor and filters, every problem, and the model when there is none. + + A field the document does not hold is `UNSET`; a `compressor` or + `filters` written as `null` is None. """ - if not isinstance(value, Mapping): - return (ValidationProblem((), "expected a mapping", "invalid_type"),) - doc = cast("Mapping[object, object]", value) - problems: list[ValidationProblem] = list(_missing_keys(ARRAY_METADATA_REQUIRED_KEYS_V3, doc)) - problems.extend(_validate_other_members(doc, ARRAY_METADATA_STANDARD_KEYS_V3)) - problems.extend(_check_literal(doc, "zarr_format", 3)) - problems.extend(_check_literal(doc, "node_type", "array")) - shape, shape_problems = _dimension_lengths(doc, "shape") + + dtype: ResolvedField[ZarrV2DataTypeDefinition[Any]] | UNSET = UNSET + """The dtype, as the scope read it.""" + compressor: ResolvedField[ZarrV2CodecDefinition[Any]] | UNSET | None = UNSET + """The compressor, as the scope read it; None when written as `null`.""" + filters: tuple[ResolvedField[ZarrV2CodecDefinition[Any]], ...] | UNSET | None = UNSET + """The filters, each as the scope read it; None when written as `null`.""" + problems: tuple[ValidationProblem, ...] = () + """Every reason the document is not a valid one.""" + metadata: ZarrV2ArrayMetadata | None = None + """The document's model, holding these fields, when there is no problem; None otherwise.""" + + def __reduce__(self) -> str | tuple[object, ...]: + # As the v3 reading: through its model, when it holds one. + if self.metadata is not None: + return (reading_of, (self.metadata,)) + return object.__reduce__(self) + + def fields(self) -> Iterator[tuple[Loc, ResolvedField[Any]]]: + """Each field the document holds, as the scope read it, where it sits: the dtype, a struct's record types after it, the compressor, each filter at its index.""" + if self.dtype is not UNSET: + yield from fields_of(self.dtype, ("dtype",)) + if self.compressor is not UNSET and self.compressor is not None: + yield from fields_of(self.compressor, ("compressor",)) + if self.filters is not UNSET and self.filters is not None: + for index, entry in enumerate(self.filters): + yield from fields_of(entry, ("filters", index)) + + +def read_array_v3( + value: object, context: Context, *, at: Loc = () +) -> tuple[ZarrV3ArrayMetadataReading, ArrayMembersV3 | None]: + """`value`, a v3 array document, as `context` read it, without its model, and its other members refined; None when it has a problem. + + `read_array_metadata_v3` builds the model from the two. A field object + in the document -- an `AcceptedField` built by hand -- is not JSON, and is refused + as such. `at` is where the document sits in the one handed in -- a document consolidated metadata holds sits three levels below + its group's -- so the levels a reader walks are counted from that + one's root; the problems are located in this document. + """ + if not is_object(value): + return ZarrV3ArrayMetadataReading(problems=not_an_object(value)), None + doc = value + problems: list[ValidationProblem] = list(missing_keys(ARRAY_METADATA_REQUIRED_KEYS_V3, doc)) + past = members_past_the_levels(doc, at) + problems.extend(past.values()) + whole, doc = doc, {key: item for key, item in doc.items() if key not in past} + extra_fields, found = other_members(doc, ARRAY_METADATA_STANDARD_KEYS_V3, at=at) + problems.extend(found) + problems.extend(check_literal(doc, "zarr_format", 3)) + problems.extend(check_literal(doc, "node_type", "array")) + shape, shape_problems = dimension_lengths(doc, "shape") problems.extend(shape_problems) # Each extension point is read by `resolve`, which judges its envelope # -- every extension *point* must be understood, so a `must_understand` @@ -390,21 +559,20 @@ def validate_array_metadata_v3( # configuration, against the definition in `context` that claims its # name. `must_understand: false` keeps its meaning where it has one: an # unknown top-level extension *field*, which a reader really can skip. - read: dict[str, Resolved[Any]] = {} + read: dict[str, ResolvedField[Any]] = {} for key, kind in _EXTENSION_POINTS_V3: if key in doc: - read[key], found = resolve(doc[key], kind, context, (key,)) - problems.extend(found) + read[key], found = resolve(doc[key], kind, context, (*at, key)) + problems.extend(within(found, at)) # The fill value is JSON, and judged by the data type the scope read, # when there is one: a data type nothing in scope claims leaves it # unjudged. + fill_value: JSONValue = None if "fill_value" in doc: - if "data_type" in read: - problems.extend( - fill_value_problems(read["data_type"], doc["fill_value"], ("fill_value",)) - ) - else: - problems.extend(_prefix("fill_value", validate_json(doc["fill_value"]))) + fill_value, found = refine_json(doc["fill_value"], (*at, "fill_value")) + problems.extend(within(found, at)) + if len(found) == 0 and "data_type" in read: + problems.extend(fill_value_problems(read["data_type"], fill_value, ("fill_value",))) # The chunk grid is judged against the shape, once both are read, and # says the lengths of the chunks the first codec is handed: an entry # for each dimension of the shape, None where nothing says it. @@ -412,26 +580,30 @@ def validate_array_metadata_v3( if "chunk_grid" in read and shape is not None: lengths, found = chunk_grid_lengths(read["chunk_grid"], shape, ("chunk_grid",)) problems.extend(found) - listed: dict[str, list[Resolved[Any]]] = {} + listed: dict[str, list[ResolvedField[Any]]] = {} for key, kind in _EXTENSION_LISTS_V3: if key in doc: entries = doc[key] if not _is_array(entries): - problems.append(ValidationProblem((key,), "expected a sequence", "invalid_type")) + problems.append(ValidationProblem((key,), "expected an array", "invalid_type")) else: listed[key] = [] for index, entry in enumerate(entries): - resolved, found = resolve(entry, kind, context, (key, index)) + resolved, found = resolve(entry, kind, context, (*at, key, index)) listed[key].append(resolved) - problems.extend(found) + problems.extend(within(found, at)) # The codecs are read as a pipeline, the first handed the grid's chunks # of the array's data type: in order, each judged against the chunk it # is handed. That holds one array -> bytes codec, so it is not empty. + chunk = Chunk(lengths, read.get("data_type")) + pipeline: tuple[Stage, ...] = () if "codecs" in listed: - chunk = Chunk(lengths, read.get("data_type")) - problems.extend(read_pipeline(listed["codecs"], chunk, ("codecs",))[1]) + pipeline, found = read_pipeline(listed["codecs"], chunk, ("codecs",)) + problems.extend(found) + attributes: dict[str, JSONValue] | None = {} if "attributes" in doc: - problems.extend(_validate_attributes(doc["attributes"])) + attributes, found = attributes_of(doc["attributes"], at) + problems.extend(found) if "dimension_names" in doc: # Simple typed sequences (dimension_names, shape, chunks) report a single # field-level loc, not per-bad-item locs; per-index locs are reserved for @@ -439,12 +611,12 @@ def validate_array_metadata_v3( names = doc["dimension_names"] if not _is_array(names): problems.append( - ValidationProblem(("dimension_names",), "expected a sequence", "invalid_type") + ValidationProblem(("dimension_names",), "expected an array", "invalid_type") ) elif not all(item is None or isinstance(item, str) for item in names): problems.append( ValidationProblem( - ("dimension_names",), "expected items of str or None", "invalid_type" + ("dimension_names",), "expected an array of strings or null", "invalid_type" ) ) elif shape is not None and len(names) != len(shape): @@ -455,53 +627,104 @@ def validate_array_metadata_v3( "invalid_value", ) ) - return tuple(problems) + reading = ZarrV3ArrayMetadataReading( + data_type=read.get("data_type", UNSET), + chunk_grid=read.get("chunk_grid", UNSET), + chunk_key_encoding=read.get("chunk_key_encoding", UNSET), + chunk=chunk, + pipeline=pipeline, + storage_transformers=tuple(listed.get("storage_transformers", ())), + problems=with_input(problems, whole), + ) + if len(problems) != 0 or shape is None or attributes is None: + return reading, None + names = doc.get("dimension_names", UNSET) + members = ArrayMembersV3( + shape=shape, + fill_value=fill_value, + dimension_names=UNSET if names is UNSET else tuple(cast("Sequence[str | None]", names)), + attributes=attributes, + extra_fields=extra_fields, + ) + return reading, members + + +def validate_array_metadata_v3( + value: object, *, context: ZarrV3Context | None = None +) -> tuple[ValidationProblem, ...]: + """Return every reason `value` is not a valid v3 array document. + + Its structure, and each extension point read through the definition + that claims its name in `context`: a gzip `level` out of range, a key a + codec's configuration does not declare. The fill value is judged + against the data type as `context` read it -- an `int8` fill value of + 300 -- and the chunk grid against the shape: a regular grid with a + chunk length for each of two dimensions, over an array of three. The + codecs are read as a pipeline: in order, each judged against the chunk + it is handed -- a `transpose` whose `order` has another number of + axes, a shard its inner chunks do not divide -- and a shard's inner + and index codecs too. + A name nothing in `context` claims is left unjudged, with any fill + value of it, and a codec of that name leaves the codec after it + handed a chunk nothing is known of. Unknown top-level keys are + allowed (they map to `extra_fields`); a reader must understand each + one that does not say `must_understand: false`, which the model + reports as `must_understand_fields`. These are the `problems` of + `read_array_metadata_v3`, which holds what was read to find them. + """ + scope = scoped(context, CORE_AND_EXTENSIONS) + return read_array_v3(value, scope)[0].problems def is_array_metadata_v3( - value: object, *, context: Context = CORE_AND_EXTENSIONS + value: object, *, context: ZarrV3Context | None = None ) -> TypeGuard[ZarrV3ArrayMetadataJSON]: """Whether `value` is a v3 array document `validate_array_metadata_v3` finds nothing wrong with, written with tuples.""" + scope = scoped(context, CORE_AND_EXTENSIONS) return ( _is_canonical_json(value, finite=False) - and not validate_array_metadata_v3(value, context=context) - and _is_canonical_array_metadata_v3(value) + and not validate_array_metadata_v3(value, context=scope) + and is_canonical_array_metadata_v3(value) ) def parse_array_metadata_v3( - value: object, *, context: Context = CORE_AND_EXTENSIONS + value: object, *, context: ZarrV3Context | None = None ) -> ZarrV3ArrayMetadataJSON: """Return `value` as `ZarrV3ArrayMetadataJSON`, or raise `MetadataValidationError`.""" - normalized = arrays_to_tuples(value) - problems = validate_array_metadata_v3(normalized, context=context) + scope = scoped(context, CORE_AND_EXTENSIONS) + problems = validate_array_metadata_v3(value, context=scope) if len(problems) != 0: raise MetadataValidationError(problems) - return cast("ZarrV3ArrayMetadataJSON", normalized) + return cast("ZarrV3ArrayMetadataJSON", arrays_to_tuples(value)) -def validate_array_metadata_v2(value: object) -> tuple[ValidationProblem, ...]: - """Return every reason `value` is not a structurally-valid v2 array doc. +def read_array_v2( + value: object, context: Context, *, at: Loc = () +) -> tuple[ZarrV2ArrayMetadataReading, ArrayMembersV2 | None]: + """`value`, a v2 array document, read once in `context`: the reading, and its other members when nothing is wrong. - Checks structure, not domain validity: `dtype` must be a string or field - records, but the string content is not interpreted; `compressor` and - `filters` are required keys that may be `None`, and otherwise must be - codec configurations (mappings with a string `id`). + `dtype`, `compressor` and `filters` are read in `context`: a dtype or + codec the scope refuses is a problem, one it does not claim is not; + `fill_value` is judged by the dtype the scope read. `compressor` and + `filters` are required keys that may be `None`. `at` is where the + document sits in the one handed in, which prefixes every problem. """ - if not isinstance(value, Mapping): - return (ValidationProblem((), "expected a mapping", "invalid_type"),) - doc = cast("Mapping[object, object]", value) + if not is_object(value): + return ZarrV2ArrayMetadataReading(problems=within(not_an_object(value), at)), None + doc = value # Unlike the group document ("Other keys MUST NOT be present", # https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L313), the v2 array document is open: other keys "SHOULD NOT be # present within the metadata object and SHOULD be ignored by # implementations" (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L91-L92), so a member outside # ARRAY_METADATA_STANDARD_KEYS_V2 is not a problem for being there. Ignored # is not unchecked: it is JSON, and its key a string, as in v3. - problems: list[ValidationProblem] = list(_missing_keys(ARRAY_METADATA_REQUIRED_KEYS_V2, doc)) - problems.extend(_validate_other_members(doc, ARRAY_METADATA_STANDARD_KEYS_V2)) - problems.extend(_check_literal(doc, "zarr_format", 2)) - shape, shape_problems = _dimension_lengths(doc, "shape") - chunks, chunks_problems = _dimension_lengths(doc, "chunks") + problems: list[ValidationProblem] = list(missing_keys(ARRAY_METADATA_REQUIRED_KEYS_V2, doc)) + extra_fields, found = other_members(doc, ARRAY_METADATA_STANDARD_KEYS_V2) + problems.extend(found) + problems.extend(check_literal(doc, "zarr_format", 2)) + shape, shape_problems = dimension_lengths(doc, "shape") + chunks, chunks_problems = dimension_lengths(doc, "chunks") problems.extend(shape_problems) problems.extend(chunks_problems) if shape is not None and chunks is not None and len(shape) != len(chunks): @@ -512,225 +735,150 @@ def validate_array_metadata_v2(value: object) -> tuple[ValidationProblem, ...]: "invalid_value", ) ) - if "dtype" in doc and not _is_dtype_v2(doc["dtype"]): - problems.append( - ValidationProblem( - ("dtype",), - "expected a v2 dtype string or a sequence of field records", - "invalid_type", - ) - ) + dtype: ResolvedField[ZarrV2DataTypeDefinition[Any]] | UNSET = UNSET + if "dtype" in doc: + # A typestr by its family, field records as a struct. + dtype, found = resolve_dtype_v2(doc["dtype"], context, ("dtype",)) + problems.extend(found) if "order" in doc and doc["order"] not in ("C", "F"): - problems.append( - ValidationProblem( - ("order",), f"expected 'C' or 'F', got {doc['order']!r}", "invalid_value" - ) - ) + problems.append(outside_of(("order",), doc["order"], ("C", "F"))) + compressor: ResolvedField[ZarrV2CodecDefinition[Any]] | UNSET | None = UNSET if "compressor" in doc: - compressor = doc["compressor"] - if compressor is not None: - problems.extend(_prefix("compressor", _validate_codec_v2(compressor))) + compressor = None + if doc["compressor"] is not None: + compressor, found = resolve_codec_v2(doc["compressor"], context, ("compressor",)) + problems.extend(found) + filters: tuple[ResolvedField[ZarrV2CodecDefinition[Any]], ...] | UNSET | None = UNSET if "filters" in doc: - filters = doc["filters"] - if filters is not None and ( - not _is_array(filters) or not all(_is_codec_v2(item) for item in filters) - ): - problems.append( - ValidationProblem( - ("filters",), - "expected null or a sequence of codec configurations with string 'id's", - "invalid_type", + filters = None + entries = doc["filters"] + if entries is not None: + if not _is_array(entries): + problems.append( + ValidationProblem( + ("filters",), + "expected null or an array of codec configurations", + "invalid_type", + ) ) - ) - elif _is_array(filters): - # "A list of JSON objects providing codec configurations, or - # null" (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L76-L79): an empty list is a list. - for index, item in enumerate(filters): - problems.extend(_prefix("filters", _prefix(index, validate_json(item)))) + else: + # "A list of JSON objects providing codec configurations, or + # null" (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L76-L79): an empty list is a list. + read: list[ResolvedField[ZarrV2CodecDefinition[Any]]] = [] + for index, item in enumerate(entries): + entry, found = resolve_codec_v2(item, context, ("filters", index)) + read.append(entry) + problems.extend(found) + filters = tuple(read) if "dimension_separator" in doc and doc["dimension_separator"] not in (".", "/"): problems.append( - ValidationProblem( - ("dimension_separator",), - f"expected '.' or '/', got {doc['dimension_separator']!r}", - "invalid_value", - ) + outside_of(("dimension_separator",), doc["dimension_separator"], (".", "/")) ) + fill_value: JSONValue = None if "fill_value" in doc: - problems.extend(_prefix("fill_value", validate_json(doc["fill_value"]))) + # JSON, judged by the dtype the scope read, when there is one: a + # dtype nothing in scope claims leaves it unjudged. + fill_value, found = refine_json(doc["fill_value"], ("fill_value",)) + problems.extend(found) + if len(found) == 0 and dtype is not UNSET: + problems.extend(fill_value_problems(dtype, fill_value, ("fill_value",))) + attributes: dict[str, JSONValue] | UNSET | None = UNSET if "attributes" in doc: - problems.extend(_validate_attributes(doc["attributes"])) - return tuple(problems) - - -def is_array_metadata_v2(value: object) -> TypeGuard[ZarrV2ArrayMetadataJSON]: - """Whether `value` is a structurally-valid v2 array metadata document.""" - return ( - _is_canonical_json(value, finite=False) - and not validate_array_metadata_v2(value) - and _is_canonical_array_metadata_v2(value) + attributes, found = attributes_of(doc["attributes"]) + problems.extend(found) + reading = ZarrV2ArrayMetadataReading( + dtype=dtype, + compressor=compressor, + filters=filters, + problems=within(with_input(problems, doc), at), ) + if len(problems) != 0 or shape is None or chunks is None or attributes is None: + return reading, None + members = ArrayMembersV2( + shape=shape, + chunks=chunks, + fill_value=fill_value, + order=cast("ZarrV2ArrayOrder", doc["order"]), + dimension_separator=cast( + "ZarrV2ArrayDimensionSeparator", doc.get("dimension_separator", ".") + ), + attributes=attributes, + extra_fields=extra_fields, + ) + return reading, members -def parse_array_metadata_v2(value: object) -> ZarrV2ArrayMetadataJSON: - """Return `value` as `ZarrV2ArrayMetadataJSON`, or raise `MetadataValidationError`.""" - normalized = arrays_to_tuples(value) - problems = validate_array_metadata_v2(normalized) - if len(problems) != 0: - raise MetadataValidationError(problems) - return cast("ZarrV2ArrayMetadataJSON", normalized) - - -def validate_consolidated_metadata_v3( - value: object, *, context: Context = CORE_AND_EXTENSIONS +def validate_array_metadata_v2( + value: object, *, context: ZarrV2Context | None = None ) -> tuple[ValidationProblem, ...]: - """Return every reason `value` is not a valid inline consolidated envelope. + """Every reason `value` is not a valid v2 array document, read in `context`, `CORE_V2` when none is given. - Locs are value-relative (the caller prefixes with `consolidated_metadata` - where appropriate). Entries recurse into the array and group document - validators, in `context`, so a validator verdict agrees with what - `ZarrV3ConsolidatedMetadata.from_json` accepts in the same scope. + `dtype`, `compressor` and `filters` are read in the scope: a dtype or + codec the scope refuses is a problem, one it does not claim is not; + `fill_value` is judged by the dtype the scope read. """ - if not isinstance(value, Mapping): - return (ValidationProblem((), "expected a mapping", "invalid_type"),) - env = cast("Mapping[object, object]", value) - problems: list[ValidationProblem] = [ - ValidationProblem((key,), "missing required key", "missing_key") - for key in ("kind", "must_understand", "metadata") - if key not in env - ] - problems.extend(_unexpected_keys(frozenset({"kind", "must_understand", "metadata"}), env)) - problems.extend(_check_literal(env, "kind", "inline")) - if "must_understand" in env and env["must_understand"] is not False: - problems.append(ValidationProblem(("must_understand",), "expected False", "invalid_value")) - if "metadata" in env: - entries = env["metadata"] - if not isinstance(entries, Mapping): - problems.append(ValidationProblem(("metadata",), "expected a mapping", "invalid_type")) - else: - for key, entry in cast("Mapping[object, object]", entries).items(): - if not isinstance(key, str): - problems.append( - ValidationProblem(("metadata",), f"non-string key {key!r}", "invalid_type") - ) - continue - entry_obj: object = entry - node_type: object = None - if isinstance(entry, Mapping): - node_type = cast("Mapping[object, object]", entry).get("node_type") - if node_type == "array": - problems.extend( - _prefix( - "metadata", - _prefix(key, validate_array_metadata_v3(entry_obj, context=context)), - ) - ) - elif node_type == "group": - problems.extend( - _prefix( - "metadata", - _prefix(key, validate_group_metadata_v3(entry_obj, context=context)), - ) - ) - else: - problems.append( - ValidationProblem( - ("metadata", key, "node_type"), - "expected 'array' or 'group'", - "invalid_value", - ) - ) - return tuple(problems) - + return read_array_v2(value, scoped(context, CORE_V2))[0].problems -def validate_group_metadata_v3( - value: object, *, context: Context = CORE_AND_EXTENSIONS -) -> tuple[ValidationProblem, ...]: - """Return every reason `value` is not a valid v3 group document. - - Unknown top-level keys are allowed (they map to `extra_fields`); a - reader must understand each one that does not say `must_understand: - false`, which the model reports as `must_understand_fields`. A - `consolidated_metadata` key, if present, is deep-validated (envelope - and entries) via `validate_consolidated_metadata_v3`, each array in it - read as `validate_array_metadata_v3` reads one, in `context`. - """ - if not isinstance(value, Mapping): - return (ValidationProblem((), "expected a mapping", "invalid_type"),) - doc = cast("Mapping[object, object]", value) - problems: list[ValidationProblem] = list(_missing_keys(GROUP_METADATA_REQUIRED_KEYS_V3, doc)) - problems.extend( - _validate_other_members( - doc, - GROUP_METADATA_STANDARD_KEYS_V3, - additional_reserved_keys=frozenset({"consolidated_metadata"}), - ) - ) - problems.extend(_check_literal(doc, "zarr_format", 3)) - problems.extend(_check_literal(doc, "node_type", "group")) - if "attributes" in doc: - problems.extend(_validate_attributes(doc["attributes"])) - if "consolidated_metadata" in doc and doc["consolidated_metadata"] is not None: - # consolidated_metadata: null (a historical zarr-python bug) is - # structurally accepted so those stores remain readable, but the model - # repairs it to absence on read and never writes it back. - problems.extend( - _prefix( - "consolidated_metadata", - validate_consolidated_metadata_v3(doc["consolidated_metadata"], context=context), - ) - ) - return tuple(problems) - -def is_group_metadata_v3( - value: object, *, context: Context = CORE_AND_EXTENSIONS -) -> TypeGuard[ZarrV3GroupMetadataJSON]: - """Whether `value` is a v3 group document `validate_group_metadata_v3` finds nothing wrong with, written with tuples.""" - return _is_canonical_json(value, finite=False) and not validate_group_metadata_v3( - value, context=context +def is_array_metadata_v2( + value: object, *, context: ZarrV2Context | None = None +) -> TypeGuard[ZarrV2ArrayMetadataJSON]: + """Whether `value` is a valid v2 array metadata document, read in `context`, `CORE_V2` when none is given.""" + return ( + _is_canonical_json(value, finite=False) + and not validate_array_metadata_v2(value, context=context) + and _is_canonical_array_metadata_v2(value) ) -def parse_group_metadata_v3( - value: object, *, context: Context = CORE_AND_EXTENSIONS -) -> ZarrV3GroupMetadataJSON: - """Return `value` narrowed to `ZarrV3GroupMetadataJSON`, or raise `MetadataValidationError`.""" - normalized = arrays_to_tuples(value) - problems = validate_group_metadata_v3(normalized, context=context) +def parse_array_metadata_v2( + value: object, *, context: ZarrV2Context | None = None +) -> ZarrV2ArrayMetadataJSON: + """`value` as `ZarrV2ArrayMetadataJSON`, read in `context`, `CORE_V2` when none is given; `MetadataValidationError` with every problem.""" + problems = validate_array_metadata_v2(value, context=context) if len(problems) != 0: raise MetadataValidationError(problems) - return cast(ZarrV3GroupMetadataJSON, normalized) + return cast("ZarrV2ArrayMetadataJSON", arrays_to_tuples(value)) -def validate_group_metadata_v2(value: object) -> tuple[ValidationProblem, ...]: +def validate_group_metadata_v2( + value: object, *, context: ZarrV2Context | None = None +) -> tuple[ValidationProblem, ...]: """Return every reason `value` is not a structurally-valid v2 group doc. Validates the in-memory merged form: the `.zgroup` fields plus an - optional `attributes` mapping folded in from `.zattrs`. + optional `attributes` mapping folded in from `.zattrs`. A group holds + no field a scope reads; `context` is taken as every v2 reader takes it. """ - if not isinstance(value, Mapping): - return (ValidationProblem((), "expected a mapping", "invalid_type"),) - doc = cast("Mapping[object, object]", value) - problems: list[ValidationProblem] = list(_missing_keys(GROUP_METADATA_REQUIRED_KEYS_V2, doc)) - problems.extend(_unexpected_keys(GROUP_METADATA_STANDARD_KEYS_V2, doc)) - problems.extend(_check_literal(doc, "zarr_format", 2)) + scoped(context, CORE_V2) + if not is_object(value): + return not_an_object(value) + doc = value + problems: list[ValidationProblem] = list(missing_keys(GROUP_METADATA_REQUIRED_KEYS_V2, doc)) + problems.extend(unexpected_keys(GROUP_METADATA_STANDARD_KEYS_V2, doc)) + problems.extend(check_literal(doc, "zarr_format", 2)) if "attributes" in doc: - problems.extend(_validate_attributes(doc["attributes"])) - return tuple(problems) + problems.extend(validate_attributes(doc["attributes"])) + return with_input(problems, doc) -def is_group_metadata_v2(value: object) -> TypeGuard[ZarrV2GroupMetadataJSON]: - """Whether `value` is a structurally-valid v2 group metadata document.""" - return _is_canonical_json(value, finite=False) and not validate_group_metadata_v2(value) +def is_group_metadata_v2( + value: object, *, context: ZarrV2Context | None = None +) -> TypeGuard[ZarrV2GroupMetadataJSON]: + """Whether `value` is a structurally-valid v2 group metadata document; `context` is taken as every v2 reader takes it.""" + return _is_canonical_json(value, finite=False) and not validate_group_metadata_v2( + value, context=context + ) -def parse_group_metadata_v2(value: object) -> ZarrV2GroupMetadataJSON: - """Return `value` narrowed to `ZarrV2GroupMetadataJSON`, or raise `MetadataValidationError`.""" - normalized = arrays_to_tuples(value) - problems = validate_group_metadata_v2(normalized) +def parse_group_metadata_v2( + value: object, *, context: ZarrV2Context | None = None +) -> ZarrV2GroupMetadataJSON: + """`value` narrowed to `ZarrV2GroupMetadataJSON`, or `MetadataValidationError`; `context` is taken as every v2 reader takes it.""" + problems = validate_group_metadata_v2(value, context=context) if len(problems) != 0: raise MetadataValidationError(problems) - return cast(ZarrV2GroupMetadataJSON, normalized) + return cast(ZarrV2GroupMetadataJSON, arrays_to_tuples(value)) StoreKey = TypeVar("StoreKey", bound=str) @@ -764,12 +912,13 @@ def load_store_json(mapping: Mapping[StoreKey, bytes], key: str) -> object: # raises `TypeError` on most else. raw = cast("object", stored[key]) if not isinstance(raw, bytes): - raise MetadataValidationError( - [ValidationProblem((key,), f"expected bytes, got {type(raw).__name__}", "invalid_type")] + refused = ValidationProblem( + (key,), f"expected bytes, got {type(raw).__name__}", "invalid_type" ) + raise MetadataValidationError(with_input((refused,), stored)) try: return json.loads(raw) - except (UnicodeDecodeError, ValueError) as exc: + except (UnicodeDecodeError, ValueError, RecursionError) as exc: raise MetadataValidationError( [ValidationProblem((key,), f"invalid JSON: {exc}", "invalid_json")] ) from exc diff --git a/packages/zarr-metadata/src/zarr_metadata/pydantic.py b/packages/zarr-metadata/src/zarr_metadata/pydantic.py index 2cb69f50d2..ba40f4a610 100644 --- a/packages/zarr-metadata/src/zarr_metadata/pydantic.py +++ b/packages/zarr-metadata/src/zarr_metadata/pydantic.py @@ -8,13 +8,28 @@ freely with non-pydantic code (equality, isinstance, nesting). Validation delegates to the library: a raw document routes through `from_json` (the single source of truth for validation and normalization, so pydantic's -field-level coercion can never bypass it), reading v3 extension points in -`CORE_AND_EXTENSIONS`, since a field type holds no scope; a reader with a -scope of its own calls `from_json(..., context=...)` itself. An existing model -instance passes through unchanged, and serialization emits the canonical -document via `to_json`. `MetadataValidationError` subclasses `ValueError`, -so a failed parse surfaces as a pydantic `ValidationError` carrying the -loc-annotated problem messages. +field-level coercion can never bypass it). A field type reads the fields +of its document in the scope pydantic's validation context holds, as +pydantic hands any validator its context: the context itself, when it is +a `Context`, which every field type of its format then reads in, the +other format's refusing it with `TypeError`; or, when it is a +mapping, its `"zarr_metadata_context"` item for the v3 field types and +its `"zarr_metadata_context_v2"` item for the v2 ones, so a model holding +both kinds of field names each format's scope; a format whose scope is +not given reads in its own, `CORE_AND_EXTENSIONS` or `CORE_V2`: + + TypeAdapter(zmp.ZarrV3ArrayMetadata).validate_python(document, context=SCOPE) + ArrayManifest.model_validate(data, context={"zarr_metadata_context": SCOPE}) + Mixed.model_validate(data, context={"zarr_metadata_context": V3, "zarr_metadata_context_v2": V2}) + +An existing model instance passes through unchanged, as pydantic's does, and +serialization emits the canonical document via `to_json`. A failed parse +surfaces as a pydantic `ValidationError` with one line error per problem, +as pydantic reports its own: its `type` the problem's `kind`, at the +problem's `loc` under the field's, with the `input` found there and the +problem's `ctx` (a message holding a ctx placeholder rides in the ctx as +`message`, as `_line_error` says). Annotate a field with these types, not the core classes, +which pydantic cannot build a schema for. A node's attributes may hold `NaN`, `Infinity` or `-Infinity`, which the models read and write as zarr-python does. Pydantic writes JSON by its own @@ -23,6 +38,12 @@ `model_config`; a `TypeAdapter` over a field type takes no `config`, so write its value with the model's `to_key_value` instead. +Pydantic's own JSON paths bound nesting below the 256 levels the readers +walk: `validate_json` refuses a document nested about 200 levels deep, +and `dump_json` one nested about 255, each with its own error. A document +that deep goes through `validate_python` on parsed JSON, and is written +with the model's `to_key_value`. + Usage: import zarr_metadata.pydantic as zmp @@ -37,11 +58,13 @@ class ArrayManifest(BaseModel): from __future__ import annotations -from typing import TYPE_CHECKING, Annotated, TypeVar +from typing import TYPE_CHECKING, Annotated, Any, Final, LiteralString, Protocol, TypeVar, cast -from pydantic import BeforeValidator, InstanceOf, PlainSerializer +from pydantic import BeforeValidator, InstanceOf, PlainSerializer, ValidationInfo +from pydantic_core import InitErrorDetails, PydanticCustomError, ValidationError from zarr_metadata import model as _model +from zarr_metadata._json import MetadataValidationError, ValidationProblem, is_object, value_at from zarr_metadata._pydantic_schema import ( ZarrV2ArrayMetadataJSON as _ZarrV2ArrayMetadataSchema, ) @@ -60,9 +83,9 @@ class ArrayManifest(BaseModel): from zarr_metadata._pydantic_schema import ( ZarrV3GroupMetadataJSON as _ZarrV3GroupMetadataSchema, ) -from zarr_metadata._pydantic_schema import ( - ZarrV3MetadataFieldJSON as _ZarrV3MetadataFieldSchema, -) +from zarr_metadata._sentinel import UNSET +from zarr_metadata.v2.definition import CORE_V2 +from zarr_metadata.v3._registry import CORE_AND_EXTENSIONS, Context, is_scope, scoped if TYPE_CHECKING: from collections.abc import Callable @@ -70,21 +93,108 @@ class ArrayManifest(BaseModel): _M = TypeVar("_M") -def _coerce_to(cls: type[_M], parse: Callable[[object], _M]) -> Callable[[object], _M]: - """A validator that passes instances of `cls` through and parses anything else.""" +def _as_pydantic_raises(cls: type[_M], value: object, read: Callable[[], _M]) -> _M: + """What `read` gives of `value`, or the `MetadataValidationError` it raises as a `pydantic_core.ValidationError`: one line error per problem, its type the problem's kind, at the problem's loc, with its input and ctx. + + A missing key's problem reports the object missing it, as pydantic's + own `missing` error does; one that holds nothing else, what it found + not being JSON a reader walks, reports what sits at its loc. + """ + try: + return read() + except MetadataValidationError as error: + line_errors = [_line_error(problem, value) for problem in error.problems] + raise ValidationError.from_exception_data(cls.__name__, line_errors) from error + + +def _line_error(problem: ValidationProblem, value: object) -> InitErrorDetails: + """`problem` as one of a `ValidationError`'s line errors: its type the problem's kind, at its loc, with its input and ctx, and its message as `msg`. + + pydantic renders a message as a template of its ctx, each `{key}` the + ctx holds replaced, one key at a time in the ctx's order, with no way + to escape one. A message that holds one -- a document value written + `{expected}`, shown in it -- is handed whole as the ctx's `message`, + last, where what it carries is scanned for no key after it, and the + template is that placeholder, so `msg` is the problem's message + whatever it holds; a ctx member of that name yields to it. + """ + template, ctx = problem.message, dict(problem.ctx) + if any(f"{{{key}}}" in template for key in ctx): + ctx.pop("message", None) + template, ctx = "{message}", {**ctx, "message": problem.message} + return InitErrorDetails( + type=PydanticCustomError(problem.kind, cast("LiteralString", template), ctx), + loc=problem.loc, + input=_input_of(problem, value), + ) + + +def _input_of(problem: ValidationProblem, value: object) -> object: + """What a line error reports as its input: what the problem found; for a key that is missing, what `value` holds at the loc above, as pydantic's own `missing` reports the object; for what the problem could not hold, not being JSON a reader walks, what sits at its loc; and `value` itself where that is nothing either.""" + if problem.input is not UNSET: + return problem.input + if problem.kind == "missing_key": + above = value_at(value, problem.loc[:-1]) if len(problem.loc) != 0 else UNSET + return value if above is UNSET else above + there = value_at(value, problem.loc) + return value if there is UNSET else there + + +CONTEXT_KEY: Final = "zarr_metadata_context" +"""The key of a mapping validation context under which the scope the v3 field types read in sits.""" +CONTEXT_KEY_V2: Final = "zarr_metadata_context_v2" +"""The key of a mapping validation context under which the scope the v2 field types read in sits.""" + + +_Read_co = TypeVar("_Read_co", covariant=True) - def coerce(value: object) -> _M: + +class _Reads(Protocol[_Read_co]): + def __call__(self, data: object, /, *, context: Context) -> _Read_co: ... + + +def _read_in_scope( + cls: type[_M], read: _Reads[_M], default: Context[Any], key: str +) -> Callable[[object, ValidationInfo], _M]: + """A validator that passes instances of `cls` through and reads anything else in the scope the validation context holds under `key`, `default` when it holds none.""" + + def coerce(value: object, info: ValidationInfo) -> _M: if isinstance(value, cls): return value - return parse(value) + return _as_pydantic_raises( + cls, value, lambda: read(value, context=_scope(info.context, default, key)) + ) return coerce +def _scope(context: object, default: Context[Any], key: str) -> Context[Any]: + """The scope a validation context holds for one format: itself, a `Context`, the scope of every field type of its format; its `key` item, the format's own; or, holding none, `default`, the format's core scope. + + A bare `Context` of the other format is a `TypeError`, as it is at + every reader: a model holding fields of both formats names each + format's scope by its key. + """ + if is_scope(context): + return scoped(context, default) + if not is_object(context) or key not in context: + return default + scope = context[key] + if not is_scope(scope): + msg = f"{key}: the scope to read in is a Context, got {scope!r}" + raise TypeError(msg) + return scope + + ZarrV3ArrayMetadata = Annotated[ InstanceOf[_model.ZarrV3ArrayMetadata], BeforeValidator( - _coerce_to(_model.ZarrV3ArrayMetadata, _model.ZarrV3ArrayMetadata.from_json), + _read_in_scope( + _model.ZarrV3ArrayMetadata, + _model.ZarrV3ArrayMetadata.from_json, + CORE_AND_EXTENSIONS, + CONTEXT_KEY, + ), json_schema_input_type=_ZarrV3ArrayMetadataSchema, ), PlainSerializer(_model.ZarrV3ArrayMetadata.to_json, return_type=_ZarrV3ArrayMetadataSchema), @@ -94,7 +204,12 @@ def coerce(value: object) -> _M: ZarrV2ArrayMetadata = Annotated[ InstanceOf[_model.ZarrV2ArrayMetadata], BeforeValidator( - _coerce_to(_model.ZarrV2ArrayMetadata, _model.ZarrV2ArrayMetadata.from_json), + _read_in_scope( + _model.ZarrV2ArrayMetadata, + _model.ZarrV2ArrayMetadata.from_json, + CORE_V2, + CONTEXT_KEY_V2, + ), json_schema_input_type=_ZarrV2ArrayMetadataSchema, ), PlainSerializer(_model.ZarrV2ArrayMetadata.to_json, return_type=_ZarrV2ArrayMetadataSchema), @@ -104,7 +219,12 @@ def coerce(value: object) -> _M: ZarrV3GroupMetadata = Annotated[ InstanceOf[_model.ZarrV3GroupMetadata], BeforeValidator( - _coerce_to(_model.ZarrV3GroupMetadata, _model.ZarrV3GroupMetadata.from_json), + _read_in_scope( + _model.ZarrV3GroupMetadata, + _model.ZarrV3GroupMetadata.from_json, + CORE_AND_EXTENSIONS, + CONTEXT_KEY, + ), json_schema_input_type=_ZarrV3GroupMetadataSchema, ), PlainSerializer(_model.ZarrV3GroupMetadata.to_json, return_type=_ZarrV3GroupMetadataSchema), @@ -114,7 +234,12 @@ def coerce(value: object) -> _M: ZarrV2GroupMetadata = Annotated[ InstanceOf[_model.ZarrV2GroupMetadata], BeforeValidator( - _coerce_to(_model.ZarrV2GroupMetadata, _model.ZarrV2GroupMetadata.from_json), + _read_in_scope( + _model.ZarrV2GroupMetadata, + _model.ZarrV2GroupMetadata.from_json, + CORE_V2, + CONTEXT_KEY_V2, + ), json_schema_input_type=_ZarrV2GroupMetadataSchema, ), PlainSerializer(_model.ZarrV2GroupMetadata.to_json, return_type=_ZarrV2GroupMetadataSchema), @@ -124,9 +249,11 @@ def coerce(value: object) -> _M: ZarrV3ConsolidatedMetadata = Annotated[ InstanceOf[_model.ZarrV3ConsolidatedMetadata], BeforeValidator( - _coerce_to( + _read_in_scope( _model.ZarrV3ConsolidatedMetadata, _model.ZarrV3ConsolidatedMetadata.from_json, + CORE_AND_EXTENSIONS, + CONTEXT_KEY, ), json_schema_input_type=_ZarrV3ConsolidatedMetadataSchema, ), @@ -140,9 +267,11 @@ def coerce(value: object) -> _M: ZarrV2ConsolidatedMetadata = Annotated[ InstanceOf[_model.ZarrV2ConsolidatedMetadata], BeforeValidator( - _coerce_to( + _read_in_scope( _model.ZarrV2ConsolidatedMetadata, _model.ZarrV2ConsolidatedMetadata.from_json, + CORE_V2, + CONTEXT_KEY_V2, ), json_schema_input_type=_ZarrV2ConsolidatedMetadataSchema, ), @@ -153,22 +282,13 @@ def coerce(value: object) -> _M: ] """Field type for a v2 `.zmetadata` document.""" -ZarrV3MetadataField = Annotated[ - InstanceOf[_model.ZarrV3NamedConfig], - BeforeValidator( - _coerce_to(_model.ZarrV3NamedConfig, _model.ZarrV3NamedConfig.from_json), - json_schema_input_type=_ZarrV3MetadataFieldSchema, - ), - PlainSerializer(_model.ZarrV3NamedConfig.to_json, return_type=_ZarrV3MetadataFieldSchema), -] -"""Field type for one normalized v3 metadata extension envelope.""" - __all__ = [ + "CONTEXT_KEY", + "CONTEXT_KEY_V2", "ZarrV2ArrayMetadata", "ZarrV2ConsolidatedMetadata", "ZarrV2GroupMetadata", "ZarrV3ArrayMetadata", "ZarrV3ConsolidatedMetadata", "ZarrV3GroupMetadata", - "ZarrV3MetadataField", ] diff --git a/packages/zarr-metadata/src/zarr_metadata/typed_json.py b/packages/zarr-metadata/src/zarr_metadata/typed_json.py index 6fb5c14db4..9ee3ef8688 100644 --- a/packages/zarr-metadata/src/zarr_metadata/typed_json.py +++ b/packages/zarr-metadata/src/zarr_metadata/typed_json.py @@ -31,10 +31,18 @@ - Each annotation is evaluated in the module of the class that wrote it, as the spec has it, where `get_type_hints` would read what a subclass inherited in the subclass's module. -- `ReadOnly` and `Annotated` are peeled wherever they are written; a - `NewType` reads as the type it names, a type alias as the type it - stands for, and a TypedDict or alias that holds itself as deep as the - value goes. +- `Required`, `NotRequired` and `ReadOnly` are peeled wherever they are + written; a `NewType` reads as the type it names, a type alias as the + type it stands for, and a TypedDict or alias that holds itself as deep + as the value goes. +- A number's type may carry bounds, in annotated-types' vocabulary as + pydantic reads it: `Gt`, `Ge`, `Lt`, `Le` and `Interval`, at any depth, + so `tuple[Annotated[int, Ge(1)], ...]` bounds each element. A value out + of them is a problem, `invalid_value`, whose message says what the type + admits and whose `ctx` holds the bounds. `Annotated` may also carry a + note, a string or a `Doc`; any other metadata is a `TypeError`, since + a constraint `check` does not read would be one it does not hold a + value to. - A union of TypedDicts that each require a key as a `Literal` of values of their own is read by the branch that key names, and its problems are that branch's. Otherwise a value is read by the first branch it @@ -47,7 +55,26 @@ exceptions: `ValidationProblem(loc, message, kind)`, with `kind` one of `missing_key`, `invalid_type`, `invalid_value`, `unknown_key` and `invalid_json`, so a caller that tolerates a key the TypedDict does not -declare can tell it from a wrong value. +declare can tell it from a wrong value. Each carries its message's data, +as pydantic's errors and zod's issues do: `input`, the JSON the value +holds at `loc`, and `ctx`, what was expected there -- a type's bounds, or +the values of a `Literal`. + +`json_schema` writes what `check` reads as a JSON Schema, draft 2020-12, +for a validator in another language, or an editor: the JSON Schema of +the values `check` finds no problem with. A TypedDict is an object of its +keys, closed or not as it says; a bound is JSON Schema's keyword for it, +`Interval(ge=0, le=9)` a `minimum` and a `maximum`; a TypedDict or a type +alias is written once, in `$defs`, under its name. JSON Schema takes a +number with no fraction, `1.0`, for an integer, where `check` wants `1`: + + from zarr_metadata.typed_json import json_schema + + json_schema(GzipCodecConfiguration) + # {'$schema': 'https://json-schema.org/draft/2020-12/schema', + # 'type': 'object', + # 'properties': {'level': {'type': 'integer', 'minimum': 0, 'maximum': 9}}, + # 'required': ['level'], 'additionalProperties': False} `check` reads the shapes JSON takes and no others -- `int`, `float` for any number, `bool`, `str`, `None`, `JSONValue`, a `Literal`, @@ -66,14 +93,19 @@ from zarr_metadata._common import JSONValue from zarr_metadata._json import ProblemKind, ValidationProblem -from zarr_metadata._typed_json import Loc, TypedDictKeys, check, typeddict_keys +from zarr_metadata._typed_json import ( + JSONSchema, + Loc, + check, + json_schema, +) __all__ = [ + "JSONSchema", "JSONValue", "Loc", "ProblemKind", - "TypedDictKeys", "ValidationProblem", "check", - "typeddict_keys", + "json_schema", ] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/codec.py b/packages/zarr-metadata/src/zarr_metadata/v2/_codec_json.py similarity index 67% rename from packages/zarr-metadata/src/zarr_metadata/v2/codec.py rename to packages/zarr-metadata/src/zarr_metadata/v2/_codec_json.py index 69125544e6..68d17a59c0 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v2/codec.py +++ b/packages/zarr-metadata/src/zarr_metadata/v2/_codec_json.py @@ -1,9 +1,4 @@ -""" -Zarr v2 codec configuration shape. - -In v2, compressors and filters are numcodecs configuration dicts: a required -`id` field naming the codec, plus arbitrary codec-specific extra fields. -""" +"""The JSON shape of a v2 codec, in a module of its own so the array document and the codec definitions can both import it.""" from typing_extensions import TypedDict @@ -24,6 +19,4 @@ class ZarrV2CodecMetadata(TypedDict, extra_items=JSONValue): id: str -__all__ = [ - "ZarrV2CodecMetadata", -] +__all__ = ["ZarrV2CodecMetadata"] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/_definition.py b/packages/zarr-metadata/src/zarr_metadata/v2/_definition.py new file mode 100644 index 0000000000..87dba74dc0 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/_definition.py @@ -0,0 +1,277 @@ +"""The kinds of field a Zarr v2 array document holds: a dtype, and a codec in `compressor` or `filters`. + +A v2 dtype is written as NumPy writes a typestr -- a byte order, a type +code and a size, `|])([A-Za-z])([0-9]*)(?:\[([0-9]*)([A-Za-z\u03bc]+)\])?") +"""A typestr: a byte order, a type code, a size in bytes if any, and a bracketed time unit with its multiplier if any.""" + + +def parse_typestr(name: str) -> tuple[str, dict[str, JSONValue]] | None: + """`name` split into its type code and what it carries; None when it is no typestr. + + What is carried is the byte order; the item size, which only `|O` + leaves out; and, when a unit is bracketed, the unit + and its multiplier, 1 when none is written. + """ + match = TYPESTR_PATTERN.fullmatch(name) + if match is None: + return None + byteorder, code, size, factor, unit = match.groups() + if size == "" and code != "O": + # "An integer specifying the number of bytes": only `|O` writes none. + return None + carried: dict[str, JSONValue] = {"byteorder": byteorder} + if size != "": + carried["itemsize"] = int(size) + if unit is not None: + carried["unit"] = unit + carried["scale_factor"] = 1 if factor == "" else int(factor) + return code, carried + + +def typestr_problem(name: str, at: Loc) -> ValidationProblem | None: + """The problem `name`, at `at`, is when it is no typestr; None when it is one, or is `struct`.""" + if name == STRUCT_NAME: + return None + if TYPESTR_PATTERN.fullmatch(name) is not None: + if parse_typestr(name) is not None: + return None + # "An integer specifying the number of bytes": a listed code without one. + return ValidationProblem( + at, + f"expected a size in bytes after the type code, '' or '|', a type code and a size in " + f"bytes, ' str | None: + if parse_typestr(self.name) is not None: + return ( + f"{self.name!r} is how a document writes one type of a family, which reads as " + "the family; define the family" + ) + # Named, not a bare `super()`: a dataclass with slots is rebuilt. + return super(ZarrV2DataTypeDefinition, self)._refusal() + + @classmethod + def name_problem(cls, name: str, at: Loc) -> ValidationProblem | None: + return typestr_problem(name, at) + + @classmethod + def spelled(cls, name: str) -> tuple[str | None, dict[str, JSONValue] | None]: + if name == STRUCT_NAME: + return name, None + if name in FAMILIES.values(): + return None, None + parsed = parse_typestr(name) + if parsed is None: + return name, None + code, carried = parsed + family = FAMILIES.get(code) + if family is None: + return name, None + return family, carried + + def carrying_name(self, configuration: Mapping[str, JSONValue]) -> str | None: + if self.name == STRUCT_NAME: + return None + code = next(code for code, family in FAMILIES.items() if family == self.name) + name = f"{configuration['byteorder']}{code}{configuration.get('itemsize', '')}" + if "unit" in configuration: + factor = configuration.get("scale_factor", 1) + name += f"[{'' if factor == 1 else factor}{configuration['unit']}]" + return name + + @classmethod + def named_configuration( + cls, value: object + ) -> tuple[str | None, Mapping[str, object] | None, Problems]: + if isinstance(value, str): + return value, None, () + if isinstance(value, (list, tuple)): + return STRUCT_NAME, {"fields": cast("JSONValue", value)}, () + return None, None, () + + @classmethod + def envelope_problems(cls, value: object) -> Problems: + if isinstance(value, str): + bad = typestr_problem(value, ()) + return () if bad is None else (bad,) + if isinstance(value, (list, tuple)): + return () + return ( + ValidationProblem( + (), + "expected a v2 dtype -- a NumPy typestr, or an array of field records -- got " + f"{shown(value)}", + "invalid_type", + ), + ) + + @classmethod + def envelope_json(cls, name: str, configuration: Mapping[str, JSONValue]) -> JSONValue: + if name == STRUCT_NAME and "fields" in configuration: + return configuration["fields"] + return name + + @classmethod + def configuration_loc(cls, loc: Loc) -> Loc: + return loc + + @classmethod + def name_loc(cls, loc: Loc) -> Loc: + return loc + + +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class ZarrV2CodecDefinition(Definition[C], kind=True, format=2): + """A v2 codec: a numcodecs id, and the TypedDict its parameters are. + + A document writes `{"id": name, **parameters}`; the definition's + configuration is the parameters, read at the field itself. + """ + + label: ClassVar[str] = "v2 codec" + + @classmethod + def name_problem(cls, name: str, at: Loc) -> ValidationProblem | None: + if len(name) != 0: + return None + return ValidationProblem(at, "expected a codec id, got ''", "invalid_value") + + @classmethod + def named_configuration( + cls, value: object + ) -> tuple[str | None, Mapping[str, object] | None, Problems]: + if not is_object(value): + return None, None, () + name = value.get("id") + if not isinstance(name, str): + return None, None, () + return ( + name, + {key: item for key, item in value.items() if isinstance(key, str) and key != "id"}, + (), + ) + + @classmethod + def envelope_problems(cls, value: object) -> Problems: + if not is_object(value): + return ( + ValidationProblem( + (), "expected a codec configuration with a string 'id'", "invalid_type" + ), + ) + entry = value + if "id" not in entry: + return (ValidationProblem(("id",), "missing required key", "missing_key"),) + if not isinstance(entry["id"], str): + return ( + ValidationProblem( + ("id",), + f"expected a string codec id, got {shown(entry['id'])}", + "invalid_type", + ), + ) + bad = cls.name_problem(entry["id"], ("id",)) + return () if bad is None else (bad,) + + @classmethod + def envelope_json(cls, name: str, configuration: Mapping[str, JSONValue]) -> JSONValue: + return {"id": name, **configuration} + + @classmethod + def configuration_loc(cls, loc: Loc) -> Loc: + return loc + + @classmethod + def name_loc(cls, loc: Loc) -> Loc: + return (*loc, "id") + + +__all__ = [ + "FAMILIES", + "STRUCT_NAME", + "TYPESTR_PATTERN", + "ZarrV2CodecDefinition", + "ZarrV2DataTypeDefinition", + "ZarrV2DataTypeField", + "parse_typestr", + "typestr_problem", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/array.py b/packages/zarr-metadata/src/zarr_metadata/v2/array.py index 1b02e1ffc7..5c0f75dce6 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v2/array.py +++ b/packages/zarr-metadata/src/zarr_metadata/v2/array.py @@ -6,7 +6,7 @@ from typing_extensions import TypeAliasType, TypedDict from zarr_metadata._common import JSONValue -from zarr_metadata.v2.codec import ZarrV2CodecMetadata +from zarr_metadata.v2._codec_json import ZarrV2CodecMetadata ZarrV2DataTypeMetadata = TypeAliasType( "ZarrV2DataTypeMetadata", diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/codec/__init__.py b/packages/zarr-metadata/src/zarr_metadata/v2/codec/__init__.py new file mode 100644 index 0000000000..71bab561a7 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/codec/__init__.py @@ -0,0 +1,81 @@ +"""Zarr v2 codecs: the configuration shape, and one definition per numcodecs id this package models. + +In v2, compressors and filters are numcodecs configuration dicts: a required +`id` field naming the codec, plus codec-specific parameters. +""" + +from typing import Any, Final + +from zarr_metadata.v2._codec_json import ZarrV2CodecMetadata +from zarr_metadata.v2._definition import ZarrV2CodecDefinition +from zarr_metadata.v2.codec.checksum import ADLER32_V2, CRC32_V2, CRC32C_V2, FLETCHER32_V2 +from zarr_metadata.v2.codec.compression import ( + BLOSC_V2, + BZ2_V2, + GZIP_V2, + LZ4_V2, + LZMA_V2, + ZLIB_V2, + ZSTD_V2, +) +from zarr_metadata.v2.codec.filters import ( + ASTYPE_V2, + BITROUND_V2, + DELTA_V2, + FIXEDSCALEOFFSET_V2, + PACKBITS_V2, + QUANTIZE_V2, + SHUFFLE_V2, +) +from zarr_metadata.v2.codec.vlen import VLEN_ARRAY_V2, VLEN_BYTES_V2, VLEN_UTF8_V2 + +V2_CODECS: Final[tuple[ZarrV2CodecDefinition[Any], ...]] = ( + ZLIB_V2, + GZIP_V2, + BZ2_V2, + LZMA_V2, + BLOSC_V2, + ZSTD_V2, + LZ4_V2, + SHUFFLE_V2, + DELTA_V2, + FIXEDSCALEOFFSET_V2, + QUANTIZE_V2, + BITROUND_V2, + ASTYPE_V2, + PACKBITS_V2, + VLEN_UTF8_V2, + VLEN_BYTES_V2, + VLEN_ARRAY_V2, + CRC32_V2, + CRC32C_V2, + ADLER32_V2, + FLETCHER32_V2, +) +"""Every codec numcodecs 0.16 configures that this package models.""" + +__all__ = [ + "ADLER32_V2", + "ASTYPE_V2", + "BITROUND_V2", + "BLOSC_V2", + "BZ2_V2", + "CRC32C_V2", + "CRC32_V2", + "DELTA_V2", + "FIXEDSCALEOFFSET_V2", + "FLETCHER32_V2", + "GZIP_V2", + "LZ4_V2", + "LZMA_V2", + "PACKBITS_V2", + "QUANTIZE_V2", + "SHUFFLE_V2", + "V2_CODECS", + "VLEN_ARRAY_V2", + "VLEN_BYTES_V2", + "VLEN_UTF8_V2", + "ZLIB_V2", + "ZSTD_V2", + "ZarrV2CodecMetadata", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/codec/_dtype.py b/packages/zarr-metadata/src/zarr_metadata/v2/codec/_dtype.py new file mode 100644 index 0000000000..9f75076529 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/codec/_dtype.py @@ -0,0 +1,50 @@ +"""A rule for the dtype parameters of a v2 codec.""" + +from __future__ import annotations + +from dataclasses import dataclass +from typing import TYPE_CHECKING, cast + +from zarr_metadata._json import ValidationProblem +from zarr_metadata.v2._definition import parse_typestr, typestr_problem + +if TYPE_CHECKING: + from collections.abc import Iterator, Mapping + + from zarr_metadata.v3._definition import Nested + + +@dataclass(frozen=True, slots=True) +class _DtypeParameter: + """The rule `dtype_parameter` gives: a value rather than a closure, so a scope holding it pickles.""" + + keys: tuple[str, ...] + float_only: bool + + def __call__( + self, configuration: Mapping[str, object], nested: Nested + ) -> Iterator[ValidationProblem]: + for key in self.keys: + if key not in configuration: + continue + written = cast("str", configuration[key]) + bad = typestr_problem(written, (key,)) + if bad is not None: + yield bad + continue + parsed = parse_typestr(written) + if self.float_only and (parsed is None or parsed[0] != "f"): + yield ValidationProblem( + (key,), f"expected a float typestr, ' _DtypeParameter: + """The rule that each of `keys`, when written, is a typestr: numcodecs writes `np.dtype(x).str` there, and a struct is no element type. + + With `float_only`, each must name a float, as `quantize` requires. + """ + return _DtypeParameter(keys, float_only) + + +__all__ = ["dtype_parameter"] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/codec/checksum.py b/packages/zarr-metadata/src/zarr_metadata/v2/codec/checksum.py new file mode 100644 index 0000000000..8337b458bb --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/codec/checksum.py @@ -0,0 +1,28 @@ +"""The v2 checksums numcodecs 0.16 configures: `crc32`, `crc32c`, `adler32` and `fletcher32`.""" + +from __future__ import annotations + +from typing import Final, Literal, NotRequired + +from typing_extensions import ReadOnly, TypedDict + +from zarr_metadata.v2._definition import ZarrV2CodecDefinition +from zarr_metadata.v3._definition import EmptyConfiguration + + +class ZarrV2Checksum32Parameters(TypedDict, closed=True): + """`numcodecs.CRC32(location=None)`, and the other 32-bit checksums: where the checksum sits, at the start by default.""" + + location: NotRequired[ReadOnly[Literal["start", "end"]]] + + +CRC32_V2: Final = ZarrV2CodecDefinition(name="crc32", configuration=ZarrV2Checksum32Parameters) +"""`numcodecs.CRC32`.""" +CRC32C_V2: Final = ZarrV2CodecDefinition(name="crc32c", configuration=ZarrV2Checksum32Parameters) +"""`numcodecs.CRC32C`.""" +ADLER32_V2: Final = ZarrV2CodecDefinition(name="adler32", configuration=ZarrV2Checksum32Parameters) +"""`numcodecs.Adler32`.""" +FLETCHER32_V2: Final = ZarrV2CodecDefinition(name="fletcher32", configuration=EmptyConfiguration) +"""`numcodecs.Fletcher32`: nothing to configure.""" + +__all__ = ["ADLER32_V2", "CRC32C_V2", "CRC32_V2", "FLETCHER32_V2", "ZarrV2Checksum32Parameters"] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/codec/compression.py b/packages/zarr-metadata/src/zarr_metadata/v2/codec/compression.py new file mode 100644 index 0000000000..5f9092ea30 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/codec/compression.py @@ -0,0 +1,109 @@ +"""The v2 compressors numcodecs 0.16 configures: `zlib`, `gzip`, `bz2`, `lzma`, `blosc`, `zstd` and `lz4`. + +Each parameter numcodecs defaults is optional; what it writes with +`get_config()` is the full object. +""" + +from __future__ import annotations + +from typing import Annotated, Final, Literal, NotRequired + +from annotated_types import Ge, Interval +from typing_extensions import ReadOnly, TypedDict + +from zarr_metadata._common import ( + JSONValue, +) +from zarr_metadata.v2._definition import ZarrV2CodecDefinition + +ZarrV2CompressionLevel = Annotated[int, Interval(ge=-1, le=9)] +"""A zlib-style compression level, 0 to 9, or -1 for zlib's default, which numcodecs writes as given.""" + + +class ZarrV2ZlibParameters(TypedDict, closed=True): + """`numcodecs.Zlib(level=1)`.""" + + level: NotRequired[ReadOnly[ZarrV2CompressionLevel]] + + +class ZarrV2GzipParameters(TypedDict, closed=True): + """`numcodecs.GZip(level=1)`.""" + + level: NotRequired[ReadOnly[ZarrV2CompressionLevel]] + + +class ZarrV2Bz2Parameters(TypedDict, closed=True): + """`numcodecs.BZ2(level=1)`: 1 to 9.""" + + level: NotRequired[ReadOnly[Annotated[int, Interval(ge=1, le=9)]]] + + +class ZarrV2LzmaParameters(TypedDict, closed=True): + """`numcodecs.LZMA(format=1, check=-1, preset=None, filters=None)`, as `lzma` takes them.""" + + format: NotRequired[ReadOnly[Annotated[int, Interval(ge=0, le=3)]]] + check: NotRequired[ReadOnly[int]] + preset: NotRequired[ReadOnly[Annotated[int, Ge(0)] | None]] + filters: NotRequired[ReadOnly[tuple[JSONValue, ...] | None]] + + +BloscCName = Literal["blosclz", "lz4", "lz4hc", "snappy", "zlib", "zstd"] +"""The compressors inside Blosc.""" + + +class ZarrV2BloscParameters(TypedDict, closed=True): + """`numcodecs.Blosc(cname='lz4', clevel=5, shuffle=1, blocksize=0, typesize=None)`: shuffle -1 is automatic, 0 none, 1 byte, 2 bit.""" + + cname: NotRequired[ReadOnly[BloscCName]] + clevel: NotRequired[ReadOnly[Annotated[int, Interval(ge=0, le=9)]]] + shuffle: NotRequired[ReadOnly[Literal[-1, 0, 1, 2]]] + blocksize: NotRequired[ReadOnly[Annotated[int, Ge(0)]]] + typesize: NotRequired[ReadOnly[Annotated[int, Ge(1)] | None]] + + +class ZarrV2ZstdParameters(TypedDict, closed=True): + """`numcodecs.Zstd(level=0, checksum=False)`: 0 is zstd's default level, and negative levels trade ratio for speed.""" + + level: NotRequired[ReadOnly[Annotated[int, Interval(ge=-131072, le=22)]]] + checksum: NotRequired[ReadOnly[bool]] + + +class ZarrV2Lz4Parameters(TypedDict, closed=True): + """`numcodecs.LZ4(acceleration=1)`.""" + + acceleration: NotRequired[ReadOnly[int]] + + +ZLIB_V2: Final = ZarrV2CodecDefinition(name="zlib", configuration=ZarrV2ZlibParameters) +"""`numcodecs.Zlib`.""" +GZIP_V2: Final = ZarrV2CodecDefinition(name="gzip", configuration=ZarrV2GzipParameters) +"""`numcodecs.GZip`.""" +BZ2_V2: Final = ZarrV2CodecDefinition(name="bz2", configuration=ZarrV2Bz2Parameters) +"""`numcodecs.BZ2`.""" +LZMA_V2: Final = ZarrV2CodecDefinition(name="lzma", configuration=ZarrV2LzmaParameters) +"""`numcodecs.LZMA`.""" +BLOSC_V2: Final = ZarrV2CodecDefinition(name="blosc", configuration=ZarrV2BloscParameters) +"""`numcodecs.Blosc`.""" +ZSTD_V2: Final = ZarrV2CodecDefinition(name="zstd", configuration=ZarrV2ZstdParameters) +"""`numcodecs.Zstd`.""" +LZ4_V2: Final = ZarrV2CodecDefinition(name="lz4", configuration=ZarrV2Lz4Parameters) +"""`numcodecs.LZ4`.""" + +__all__ = [ + "BLOSC_V2", + "BZ2_V2", + "GZIP_V2", + "LZ4_V2", + "LZMA_V2", + "ZLIB_V2", + "ZSTD_V2", + "BloscCName", + "ZarrV2BloscParameters", + "ZarrV2Bz2Parameters", + "ZarrV2CompressionLevel", + "ZarrV2GzipParameters", + "ZarrV2Lz4Parameters", + "ZarrV2LzmaParameters", + "ZarrV2ZlibParameters", + "ZarrV2ZstdParameters", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/codec/filters.py b/packages/zarr-metadata/src/zarr_metadata/v2/codec/filters.py new file mode 100644 index 0000000000..70e1c45cc7 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/codec/filters.py @@ -0,0 +1,105 @@ +"""The v2 filters numcodecs 0.16 configures: `shuffle`, `delta`, `fixedscaleoffset`, `quantize`, `bitround`, `astype` and `packbits`. + +A dtype parameter is a NumPy typestr, as numcodecs writes one; a +parameter numcodecs defaults is optional. +""" + +from __future__ import annotations + +from typing import Annotated, Final, NotRequired + +from annotated_types import Ge +from typing_extensions import ReadOnly, TypedDict + +from zarr_metadata.v2._definition import ZarrV2CodecDefinition +from zarr_metadata.v2.codec._dtype import dtype_parameter +from zarr_metadata.v3._definition import EmptyConfiguration + + +class ZarrV2ShuffleParameters(TypedDict, closed=True): + """`numcodecs.Shuffle(elementsize=4)`.""" + + elementsize: NotRequired[ReadOnly[Annotated[int, Ge(1)]]] + + +class ZarrV2DeltaParameters(TypedDict, closed=True): + """`numcodecs.Delta(dtype, astype=None)`: `astype` is `dtype` when left out.""" + + dtype: ReadOnly[str] + astype: NotRequired[ReadOnly[str]] + + +class ZarrV2FixedScaleOffsetParameters(TypedDict, closed=True): + """`numcodecs.FixedScaleOffset(offset, scale, dtype, astype=None)`.""" + + offset: ReadOnly[float | int] + scale: ReadOnly[float | int] + dtype: ReadOnly[str] + astype: NotRequired[ReadOnly[str]] + + +class ZarrV2QuantizeParameters(TypedDict, closed=True): + """`numcodecs.Quantize(digits, dtype, astype=None)`: `dtype` names a float.""" + + digits: ReadOnly[Annotated[int, Ge(0)]] + dtype: ReadOnly[str] + astype: NotRequired[ReadOnly[str]] + + +class ZarrV2BitRoundParameters(TypedDict, closed=True): + """`numcodecs.BitRound(keepbits)`.""" + + keepbits: ReadOnly[Annotated[int, Ge(0)]] + + +class ZarrV2AsTypeParameters(TypedDict, closed=True): + """`numcodecs.AsType(encode_dtype, decode_dtype)`.""" + + encode_dtype: ReadOnly[str] + decode_dtype: ReadOnly[str] + + +SHUFFLE_V2: Final = ZarrV2CodecDefinition(name="shuffle", configuration=ZarrV2ShuffleParameters) +"""`numcodecs.Shuffle`.""" +DELTA_V2: Final = ZarrV2CodecDefinition( + name="delta", configuration=ZarrV2DeltaParameters, rules=dtype_parameter("dtype", "astype") +) +"""`numcodecs.Delta`.""" +FIXEDSCALEOFFSET_V2: Final = ZarrV2CodecDefinition( + name="fixedscaleoffset", + configuration=ZarrV2FixedScaleOffsetParameters, + rules=dtype_parameter("dtype", "astype"), +) +"""`numcodecs.FixedScaleOffset`.""" +QUANTIZE_V2: Final = ZarrV2CodecDefinition( + name="quantize", + configuration=ZarrV2QuantizeParameters, + rules=dtype_parameter("dtype", "astype", float_only=True), +) +"""`numcodecs.Quantize`.""" +BITROUND_V2: Final = ZarrV2CodecDefinition(name="bitround", configuration=ZarrV2BitRoundParameters) +"""`numcodecs.BitRound`.""" +ASTYPE_V2: Final = ZarrV2CodecDefinition( + name="astype", + configuration=ZarrV2AsTypeParameters, + rules=dtype_parameter("encode_dtype", "decode_dtype"), +) +"""`numcodecs.AsType`.""" +PACKBITS_V2: Final = ZarrV2CodecDefinition(name="packbits", configuration=EmptyConfiguration) +"""`numcodecs.PackBits`: nothing to configure.""" + +__all__ = [ + "ASTYPE_V2", + "BITROUND_V2", + "DELTA_V2", + "FIXEDSCALEOFFSET_V2", + "PACKBITS_V2", + "QUANTIZE_V2", + "SHUFFLE_V2", + "ZarrV2AsTypeParameters", + "ZarrV2BitRoundParameters", + "ZarrV2DeltaParameters", + "ZarrV2FixedScaleOffsetParameters", + "ZarrV2QuantizeParameters", + "ZarrV2ShuffleParameters", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/codec/vlen.py b/packages/zarr-metadata/src/zarr_metadata/v2/codec/vlen.py new file mode 100644 index 0000000000..18c2c87d8d --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/codec/vlen.py @@ -0,0 +1,29 @@ +"""The v2 variable-length filters numcodecs 0.16 configures: `vlen-utf8`, `vlen-bytes` and `vlen-array`.""" + +from __future__ import annotations + +from typing import Final + +from typing_extensions import ReadOnly, TypedDict + +from zarr_metadata.v2._definition import ZarrV2CodecDefinition +from zarr_metadata.v2.codec._dtype import dtype_parameter +from zarr_metadata.v3._definition import EmptyConfiguration + + +class ZarrV2VLenArrayParameters(TypedDict, closed=True): + """`numcodecs.VLenArray(dtype)`: the element type of each array.""" + + dtype: ReadOnly[str] + + +VLEN_UTF8_V2: Final = ZarrV2CodecDefinition(name="vlen-utf8", configuration=EmptyConfiguration) +"""`numcodecs.VLenUTF8`: nothing to configure.""" +VLEN_BYTES_V2: Final = ZarrV2CodecDefinition(name="vlen-bytes", configuration=EmptyConfiguration) +"""`numcodecs.VLenBytes`: nothing to configure.""" +VLEN_ARRAY_V2: Final = ZarrV2CodecDefinition( + name="vlen-array", configuration=ZarrV2VLenArrayParameters, rules=dtype_parameter("dtype") +) +"""`numcodecs.VLenArray`.""" + +__all__ = ["VLEN_ARRAY_V2", "VLEN_BYTES_V2", "VLEN_UTF8_V2", "ZarrV2VLenArrayParameters"] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/data_type/__init__.py b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/__init__.py new file mode 100644 index 0000000000..85441a2828 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/__init__.py @@ -0,0 +1,42 @@ +"""The data types zarr-python 2.x writes, one definition per NumPy family.""" + +from typing import Any, Final + +from zarr_metadata.v2._definition import ZarrV2DataTypeDefinition +from zarr_metadata.v2.data_type.fixed_width import BYTES_V2, STR_V2, VOID_V2 +from zarr_metadata.v2.data_type.object import OBJECT_V2 +from zarr_metadata.v2.data_type.scalar import BOOL_V2, COMPLEX_V2, FLOAT_V2, INT_V2, UINT_V2 +from zarr_metadata.v2.data_type.struct import STRUCT_V2 +from zarr_metadata.v2.data_type.time import DATETIME64_V2, TIMEDELTA64_V2 + +V2_DATA_TYPES: Final[tuple[ZarrV2DataTypeDefinition[Any], ...]] = ( + BOOL_V2, + INT_V2, + UINT_V2, + FLOAT_V2, + COMPLEX_V2, + BYTES_V2, + STR_V2, + VOID_V2, + DATETIME64_V2, + TIMEDELTA64_V2, + OBJECT_V2, + STRUCT_V2, +) +"""Every v2 data type, by family.""" + +__all__ = [ + "BOOL_V2", + "BYTES_V2", + "COMPLEX_V2", + "DATETIME64_V2", + "FLOAT_V2", + "INT_V2", + "OBJECT_V2", + "STRUCT_V2", + "STR_V2", + "TIMEDELTA64_V2", + "UINT_V2", + "V2_DATA_TYPES", + "VOID_V2", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/data_type/fixed_width.py b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/fixed_width.py new file mode 100644 index 0000000000..d1a1d0273f --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/fixed_width.py @@ -0,0 +1,103 @@ +"""The v2 families of a fixed width per item: `bytes` (`S`), `str` (`U`) and `void` (`V`), of any size.""" + +from __future__ import annotations + +import base64 +from typing import TYPE_CHECKING, Final + +from zarr_metadata._json import ValidationProblem, shown +from zarr_metadata.v2._definition import ZarrV2DataTypeDefinition +from zarr_metadata.v2.data_type.scalar import ZarrV2ScalarConfiguration, orderless_at, sized +from zarr_metadata.v3.data_type.bytes import base64_bytes + +if TYPE_CHECKING: + from collections.abc import Iterator + + from zarr_metadata.v3._definition import Nested + +ZarrV2Base64FillValue = str | None +"""A fill value the v2 spec encodes as base64: for a fixed-length byte string, void, or a structured type (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L191-L193); or null.""" + + +def base64_fill_value_rules( + configuration: object, nested: Nested, value: ZarrV2Base64FillValue +) -> Iterator[ValidationProblem]: + """A string of standard-alphabet base64, when it is not null.""" + yield from sized_base64_rules(value, None, exact=True) + + +def sized_base64_rules( + value: ZarrV2Base64FillValue, size: int | None, *, exact: bool +) -> Iterator[ValidationProblem]: + """A string of standard-alphabet base64 of `size` bytes, or of at most `size` when not `exact`, when it is not null; any size when `size` is None, which no item of a known size gives.""" + if value is None: + return + try: + base64_bytes(value) + except ValueError: + yield ValidationProblem( + (), f"expected standard-alphabet base64, got {shown(value)}", "invalid_value" + ) + return + if size is None: + return + held = base64.b64decode(value) + if (len(held) != size) if exact else (len(held) > size): + bound = "" if exact else "at most " + yield ValidationProblem( + (), + f"expected base64 of {bound}{size} bytes, the item's size, got {len(held)} bytes", + "invalid_value", + ) + + +def _void_fill_value_rules( + configuration: ZarrV2ScalarConfiguration, nested: Nested, value: ZarrV2Base64FillValue +) -> Iterator[ValidationProblem]: + """Base64 of exactly the item's bytes.""" + yield from sized_base64_rules(value, configuration["itemsize"], exact=True) + + +def _bytes_fill_value_rules( + configuration: ZarrV2ScalarConfiguration, nested: Nested, value: ZarrV2Base64FillValue +) -> Iterator[ValidationProblem]: + """Base64 of at most the item's bytes: NumPy pads a shorter byte string with zeros.""" + yield from sized_base64_rules(value, configuration["itemsize"], exact=False) + + +BYTES_V2: Final = ZarrV2DataTypeDefinition( + name="bytes", + configuration=ZarrV2ScalarConfiguration, + rules=sized(None, None), + canonical=orderless_at(None), + fill_value=ZarrV2Base64FillValue, + fill_value_rules=_bytes_fill_value_rules, +) +"""`|S`: byte strings of `n` bytes; the fill value base64 of at most `n`.""" + +STR_V2: Final = ZarrV2DataTypeDefinition( + name="str", + configuration=ZarrV2ScalarConfiguration, + rules=sized(None, frozenset()), + fill_value=str | None, +) +"""``: strings of `n` code points, each four bytes in the byte order written; the fill value a string.""" + +VOID_V2: Final = ZarrV2DataTypeDefinition( + name="void", + configuration=ZarrV2ScalarConfiguration, + rules=sized(None, None), + canonical=orderless_at(None), + fill_value=ZarrV2Base64FillValue, + fill_value_rules=_void_fill_value_rules, +) +"""`|V`: `n` bytes of no type; the fill value base64 of exactly `n`.""" + +__all__ = [ + "BYTES_V2", + "STR_V2", + "VOID_V2", + "ZarrV2Base64FillValue", + "base64_fill_value_rules", + "sized_base64_rules", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/data_type/object.py b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/object.py new file mode 100644 index 0000000000..2ea6547cf9 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/object.py @@ -0,0 +1,31 @@ +"""The v2 `object` family, `|O`: Python objects, whose fill value is any JSON.""" + +from __future__ import annotations + +from typing import Final, cast + +from typing_extensions import ReadOnly, TypedDict + +from zarr_metadata.v2._definition import ZarrV2DataTypeDefinition +from zarr_metadata.v2.data_type.scalar import ( + ZarrV2ByteOrder, +) + + +class ZarrV2ObjectConfiguration(TypedDict, closed=True): + """What `|O` carries: a byte order, which the type ignores.""" + + byteorder: ReadOnly[ZarrV2ByteOrder] + + +def _canonical(configuration: ZarrV2ObjectConfiguration) -> ZarrV2ObjectConfiguration: + """No byte order, `|`, as NumPy writes it.""" + return cast("ZarrV2ObjectConfiguration", {**configuration, "byteorder": "|"}) + + +OBJECT_V2: Final = ZarrV2DataTypeDefinition( + name="object", configuration=ZarrV2ObjectConfiguration, canonical=_canonical +) +"""`|O`: Python objects, each encoded by a filter; the fill value any JSON.""" + +__all__ = ["OBJECT_V2", "ZarrV2ObjectConfiguration"] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/data_type/scalar.py b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/scalar.py new file mode 100644 index 0000000000..0dc2d593ab --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/scalar.py @@ -0,0 +1,234 @@ +"""The v2 scalar families: `bool`, `int`, `uint`, `float` and `complex`, filed by family and read for every typestr of the family.""" + +from __future__ import annotations + +import math +import struct +from dataclasses import dataclass +from typing import TYPE_CHECKING, Annotated, Final, Literal, cast + +from annotated_types import Ge +from typing_extensions import ReadOnly, TypedDict + +from zarr_metadata._json import ValidationProblem +from zarr_metadata.v2._definition import ZarrV2DataTypeDefinition + +if TYPE_CHECKING: + from collections.abc import Iterator, Mapping + + from zarr_metadata.v3._definition import Nested + +ZarrV2ByteOrder = Literal["<", ">", "|"] +"""The byte orders a typestr writes: little-endian, big-endian, and not relevant.""" + + +class ZarrV2ScalarConfiguration(TypedDict, closed=True): + """What a scalar typestr carries: its byte order and its size in bytes, `{"byteorder": "<", "itemsize": 4}` for ` Iterator[ValidationProblem]: + size = configuration["itemsize"] + if self.sizes is not None and size not in self.sizes: + yield ValidationProblem( + ("itemsize",), + f"expected a size of {', '.join(map(str, self.sizes))} bytes, got {size}", + "invalid_value", + ) + return + if ( + configuration["byteorder"] == "|" + and self.orderless is not None + and size not in self.orderless + ): + yield ValidationProblem( + ("byteorder",), + f"expected a byte order '<' or '>' for a type of {size} bytes, got '|'", + "invalid_value", + ) + + +def sized(sizes: tuple[int, ...] | None, orderless: frozenset[int] | None) -> _Sized: + """The rules of a family whose types are `sizes` bytes wide, any width when None. + + `orderless` is the sizes at which the type has no byte order, which + NumPy writes as `|`: one byte for an integer, every size for bytes + and void, none for a float; None is every size. At any other size + the typestr says which end comes first, `<` or `>`. + """ + return _Sized(sizes, orderless) + + +@dataclass(frozen=True, slots=True) +class _OrderlessAt: + """The canonical spelling of a family whose types of `sizes` bytes have no byte order.""" + + sizes: frozenset[int] | None + + def __call__(self, configuration: ZarrV2ScalarConfiguration) -> ZarrV2ScalarConfiguration: + orderless = self.sizes is None or configuration["itemsize"] in self.sizes + if orderless and configuration["byteorder"] != "|": + return cast("ZarrV2ScalarConfiguration", {**configuration, "byteorder": "|"}) + return configuration + + +def orderless_at(sizes: frozenset[int] | None) -> _OrderlessAt: + """The canonical spelling of a family whose types of `sizes` bytes have no byte order: `|`, as NumPy writes it; None is every size.""" + return _OrderlessAt(sizes) + + +_ONE_BYTE: Final = frozenset({1}) + + +@dataclass(frozen=True, slots=True) +class _InRange: + """The rule that an integer fill value lies in the range of the type's size.""" + + signed: bool + + def __call__( + self, configuration: ZarrV2ScalarConfiguration, nested: Nested, value: int | None + ) -> Iterator[ValidationProblem]: + if value is None: + return + bits = 8 * configuration["itemsize"] + if self.signed: + low, high = -(2 ** (bits - 1)), 2 ** (bits - 1) - 1 + else: + low, high = 0, 2**bits - 1 + if not low <= value <= high: + yield ValidationProblem( + (), f"expected an integer in [{low}, {high}], got {value}", "invalid_value" + ) + + +ZarrV2FloatSpecial = Literal["NaN", "Infinity", "-Infinity"] +"""The non-finite values the v2 spec spells by name (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L178-L190).""" + +ZarrV2FloatFillValue = float | int | ZarrV2FloatSpecial | None +"""A v2 float fill value: a number, a named non-finite value, or null.""" + +ZarrV2ComplexComponent = float | int | ZarrV2FloatSpecial +"""One component of a complex fill value: a float fill value that is not null.""" + +ZarrV2ComplexFillValue = tuple[ZarrV2ComplexComponent, ZarrV2ComplexComponent] | None +"""A v2 complex fill value: `[real, imag]`, each a float fill value, or null.""" + + +def _float_canonical( + configuration: ZarrV2ScalarConfiguration, nested: Nested, value: ZarrV2FloatFillValue +) -> ZarrV2FloatFillValue: + """An integer written for a float is the float, `0` and `0.0` one value, and a number past the largest float of the dtype's width is the infinity of its sign, as the v3 float types read one and NumPy stores it.""" + return _held_in(value, configuration["itemsize"]) + + +def _held_in(value: ZarrV2FloatFillValue, itemsize: int) -> ZarrV2FloatFillValue: + """`value` as a float of `itemsize` bytes holds it: itself as a float, or the infinity of its sign past the largest such float; a width no float has keeps the value.""" + if isinstance(value, bool) or not isinstance(value, (int, float)): + return value + try: + held = float(value) + narrowed = held + if itemsize in _FLOAT_CODES: + # `struct` refuses a float16 past the largest, and packs a + # float32 past the largest as an infinity. + code = _FLOAT_CODES[itemsize] + (narrowed,) = struct.unpack(code, struct.pack(code, held)) + except OverflowError: + return "-Infinity" if value < 0 else "Infinity" + if math.isinf(narrowed): + return "-Infinity" if value < 0 else "Infinity" + return held + + +_FLOAT_CODES: Final[Mapping[int, str]] = {2: "e", 4: "f", 8: "d"} +"""The `struct` code of the float each size in bytes is: what refuses a value past the largest.""" + + +def _complex_canonical( + configuration: ZarrV2ScalarConfiguration, nested: Nested, value: ZarrV2ComplexFillValue +) -> ZarrV2ComplexFillValue: + """Each component in the float's canonical spelling.""" + if value is None: + return None + real, imag = value + width = configuration["itemsize"] // 2 + return ( + cast("ZarrV2ComplexComponent", _held_in(real, width)), + cast("ZarrV2ComplexComponent", _held_in(imag, width)), + ) + + +BOOL_V2: Final = ZarrV2DataTypeDefinition( + name="bool", + configuration=ZarrV2ScalarConfiguration, + rules=sized((1,), None), + canonical=orderless_at(None), + fill_value=bool | None, +) +"""`|b1`: one byte, true or false.""" + +INT_V2: Final = ZarrV2DataTypeDefinition( + name="int", + configuration=ZarrV2ScalarConfiguration, + rules=sized((1, 2, 4, 8), _ONE_BYTE), + canonical=orderless_at(_ONE_BYTE), + fill_value=int | None, + fill_value_rules=_InRange(True), +) +"""`i1` to `i8`: signed integers, the fill value in the type's range.""" + +UINT_V2: Final = ZarrV2DataTypeDefinition( + name="uint", + configuration=ZarrV2ScalarConfiguration, + rules=sized((1, 2, 4, 8), _ONE_BYTE), + canonical=orderless_at(_ONE_BYTE), + fill_value=int | None, + fill_value_rules=_InRange(False), +) +"""`u1` to `u8`: unsigned integers, the fill value in the type's range.""" + +FLOAT_V2: Final = ZarrV2DataTypeDefinition( + name="float", + configuration=ZarrV2ScalarConfiguration, + rules=sized((2, 4, 8), frozenset()), + fill_value=ZarrV2FloatFillValue, + fill_value_canonical=_float_canonical, +) +"""`f2`, `f4`, `f8`: IEEE 754 floats; the fill value a number or a named non-finite value.""" + +COMPLEX_V2: Final = ZarrV2DataTypeDefinition( + name="complex", + configuration=ZarrV2ScalarConfiguration, + rules=sized((8, 16), frozenset()), + fill_value=ZarrV2ComplexFillValue, + fill_value_canonical=_complex_canonical, +) +"""`c8`, `c16`: complex floats; the fill value `[real, imag]`.""" + +__all__ = [ + "BOOL_V2", + "COMPLEX_V2", + "FLOAT_V2", + "INT_V2", + "UINT_V2", + "ZarrV2ByteOrder", + "ZarrV2ComplexComponent", + "ZarrV2ComplexFillValue", + "ZarrV2FloatFillValue", + "ZarrV2FloatSpecial", + "ZarrV2ScalarConfiguration", + "orderless_at", + "sized", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/data_type/struct.py b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/struct.py new file mode 100644 index 0000000000..979e6eb474 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/struct.py @@ -0,0 +1,94 @@ +"""The v2 `struct` type: an array of field records, each record's type a field of its own.""" + +from __future__ import annotations + +import math +from typing import TYPE_CHECKING, Annotated, Any, Final, cast + +from annotated_types import Ge +from typing_extensions import ReadOnly, TypedDict + +from zarr_metadata._json import ValidationProblem +from zarr_metadata.v2._definition import STRUCT_NAME, ZarrV2DataTypeDefinition, ZarrV2DataTypeField +from zarr_metadata.v2.data_type.fixed_width import ZarrV2Base64FillValue, sized_base64_rules +from zarr_metadata.v3._definition import AcceptedField + +if TYPE_CHECKING: + from collections.abc import Iterator + + from zarr_metadata.v3._definition import Nested + +# The longer record first: a record that fits neither is reported by the +# first branch, and a shape that is wrong is deeper than a length that is. +ZarrV2StructRecord = ( + tuple[str, ZarrV2DataTypeField, tuple[Annotated[int, Ge(0)], ...]] + | tuple[str, ZarrV2DataTypeField] +) +"""A field record: `[name, dtype]` or `[name, dtype, shape]`, the dtype a nested field (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L152-L174).""" + + +class ZarrV2StructConfiguration(TypedDict, closed=True): + """The records of a structured dtype, which the document writes as the dtype itself.""" + + fields: ReadOnly[tuple[ZarrV2StructRecord, ...]] + + +def _rules(configuration: ZarrV2StructConfiguration, nested: Nested) -> Iterator[ValidationProblem]: + """At least one record, the names distinct; `""` is NumPy's name for the padding of an aligned struct, which zarr-python 2.x writes, and may repeat.""" + fields = configuration["fields"] + if len(fields) == 0: + yield ValidationProblem(("fields",), "expected at least one field record", "invalid_value") + seen: dict[str, int] = {} + for index, record in enumerate(fields): + name = record[0] + if name == "": + continue + first = seen.setdefault(name, index) + if first != index: + yield ValidationProblem( + ("fields", index, 0), + f"duplicate field name {name!r}, already used by record {first}", + "invalid_value", + ) + + +def _fill_value_rules( + configuration: ZarrV2StructConfiguration, nested: Nested, value: ZarrV2Base64FillValue +) -> Iterator[ValidationProblem]: + """Base64 of exactly one record, when every field's size is known.""" + yield from sized_base64_rules(value, record_size(configuration, nested), exact=True) + + +def record_size(configuration: ZarrV2StructConfiguration, nested: Nested) -> int | None: + """The bytes one record of `configuration` takes: each field's item size times the product of its subarray shape, a nested struct's its own record; None when a field's size is not known, as `|O`'s is not, or its dtype was not read.""" + total = 0 + for index, record in enumerate(configuration["fields"]): + field = nested.get(("fields", index, 1)) + if not isinstance(field, AcceptedField): + return None + size = _item_size(field) + if size is None: + return None + shape = record[2] if len(record) == 3 else () + total += size * math.prod(shape) + return total + + +def _item_size(field: AcceptedField[Any]) -> int | None: + """The bytes one item of the dtype `field` read takes; None when not known.""" + if field.name == STRUCT_NAME or "fields" in field.configuration: + return record_size(cast("ZarrV2StructConfiguration", field.configuration), field.nested) + size = field.configuration.get("itemsize") + return size if isinstance(size, int) and not isinstance(size, bool) else None + + +STRUCT_V2: Final = ZarrV2DataTypeDefinition( + name=STRUCT_NAME, + configuration=ZarrV2StructConfiguration, + rules=_rules, + fill_value=ZarrV2Base64FillValue, + fill_value_rules=_fill_value_rules, +) +"""A structured type: an array of field records; the fill value base64 of one record (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L191-L193).""" + +__all__ = ["STRUCT_V2", "ZarrV2StructConfiguration", "ZarrV2StructRecord"] diff --git a/packages/zarr-metadata/src/zarr_metadata/v2/data_type/time.py b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/time.py new file mode 100644 index 0000000000..71410bfb47 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v2/data_type/time.py @@ -0,0 +1,117 @@ +"""The v2 time families, `datetime64` (`M`) and `timedelta64` (`m`): eight bytes, in a unit the typestr brackets.""" + +from __future__ import annotations + +from typing import TYPE_CHECKING, Annotated, Final, Literal, NotRequired, cast + +from annotated_types import Ge +from typing_extensions import ReadOnly, TypedDict + +from zarr_metadata._json import ValidationProblem +from zarr_metadata.v2._definition import ZarrV2DataTypeDefinition +from zarr_metadata.v2.data_type.scalar import ( + ZarrV2ByteOrder, +) +from zarr_metadata.v3.data_type._numpy_time import ( + NumpyTimeScaleFactor, + NumpyTimeTicks, + numpy_time_fill_value_canonical, +) + +if TYPE_CHECKING: + from collections.abc import Iterator + + from zarr_metadata.v3._definition import Nested + + +class ZarrV2TimeConfiguration(TypedDict, closed=True): + """What a time typestr carries: ` Iterator[ValidationProblem]: + """Eight bytes, in an order; and a unit, which the v2 spec requires (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L147-L150).""" + size = configuration["itemsize"] + if size != 8: + yield ValidationProblem( + ("itemsize",), f"expected a size of 8 bytes, got {size}", "invalid_value" + ) + return + if configuration["byteorder"] == "|": + yield ValidationProblem( + ("byteorder",), + "expected a byte order '<' or '>' for a type of 8 bytes, got '|'", + "invalid_value", + ) + if "unit" not in configuration: + yield ValidationProblem( + ("unit",), + "expected a time unit in brackets, ' ZarrV2TimeConfiguration: + """`μs` is written `us`, as NumPy writes it.""" + if configuration.get("unit") == "μs": + return cast("ZarrV2TimeConfiguration", {**configuration, "unit": "us"}) + return configuration + + +ZarrV2TimeFillValue = NumpyTimeTicks | Literal["NaT"] | None +"""A v2 time fill value: a count of ticks, `"NaT"`, or null.""" + + +def _fill_value_canonical( + configuration: ZarrV2TimeConfiguration, nested: Nested, value: ZarrV2TimeFillValue +) -> ZarrV2TimeFillValue: + """`NaT` for not-a-time however it is written; any other value as written.""" + if value is None: + return None + return cast( + "ZarrV2TimeFillValue", numpy_time_fill_value_canonical(configuration, nested, value) + ) + + +DATETIME64_V2: Final = ZarrV2DataTypeDefinition( + name="datetime64", + configuration=ZarrV2TimeConfiguration, + rules=_rules, + canonical=_canonical, + fill_value=ZarrV2TimeFillValue, + fill_value_canonical=_fill_value_canonical, +) +"""` tuple[ResolvedField[ZarrV2DataTypeDefinition[Any]], Problems]: + """`value`, a v2 `dtype`, read in `context`, `CORE_V2` when none is given: what the scope made of it, and every problem, each prefixed with `loc`.""" + return resolve(value, ZarrV2DataTypeDefinition, scoped(context, CORE_V2), loc) + + +def resolve_codec_v2( + value: object, context: ZarrV2Context | None = None, loc: Loc = () +) -> tuple[ResolvedField[ZarrV2CodecDefinition[Any]], Problems]: + """`value`, a v2 `compressor` or one of its `filters`, read in `context`, `CORE_V2` when none is given: what the scope made of it, and every problem, each prefixed with `loc`.""" + return resolve(value, ZarrV2CodecDefinition, scoped(context, CORE_V2), loc) + + +__all__ = [ + "CORE_V2", + "V2_CODECS", + "V2_DATA_TYPES", + "AcceptedField", + "ClaimKey", + "Claims", + "Conflict", + "Context", + "Disagreements", + "RefusedField", + "ResolvedField", + "ScopeConflictError", + "UnclaimedField", + "ZarrV2CodecDefinition", + "ZarrV2Context", + "ZarrV2DataTypeDefinition", + "ZarrV2DataTypeField", + "canonical_fill_value", + "fill_value_problems", + "resolve_codec_v2", + "resolve_dtype_v2", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/__init__.py b/packages/zarr-metadata/src/zarr_metadata/v3/__init__.py index 4e335f9573..9879db759f 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/__init__.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/__init__.py @@ -1,11 +1,14 @@ """Zarr v3 metadata types.""" from zarr_metadata.v3._common import ZarrV3MetadataFieldJSON +from zarr_metadata.v3._hierarchy import NodeName, NodePath from zarr_metadata.v3.array import ZarrV3ArrayMetadataJSON, ZarrV3ExtensionField from zarr_metadata.v3.consolidated import ZarrV3ConsolidatedMetadataJSON from zarr_metadata.v3.group import ZarrV3GroupMetadataJSON __all__ = [ + "NodeName", + "NodePath", "ZarrV3ArrayMetadataJSON", "ZarrV3ConsolidatedMetadataJSON", "ZarrV3ExtensionField", diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/_common.py b/packages/zarr-metadata/src/zarr_metadata/v3/_common.py index e6d76d8bea..f71b347d5c 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/_common.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/_common.py @@ -1,13 +1,16 @@ -"""The v3 metadata field: its JSON, and the validators that judge one on its own. +"""The v3 metadata field: its JSON, the aliases a member holding one is annotated with, and the validators that judge one on its own. Private, and below both readers of a field: the model, which judges the fields of a document, and the definitions, which read a field's configuration. Public consumers import `ZarrV3MetadataFieldJSON` from -`zarr_metadata.v3`, and the validators from `zarr_metadata.model`. +`zarr_metadata.v3`, the aliases from `zarr_metadata.v3.definition`, and +the validators from `zarr_metadata.model`. """ -from collections.abc import Mapping -from typing import TypeGuard, cast +import re +from typing import Final, TypeAlias, TypeGuard, cast + +from typing_extensions import TypeAliasType from zarr_metadata._common import ZarrV3NamedConfigJSON from zarr_metadata._json import ( @@ -15,11 +18,14 @@ ValidationProblem, arrays_to_tuples, is_canonical_json, - prefixed, + is_object, + shown, + shown_key, validate_json, + with_input, ) -ZarrV3MetadataFieldJSON = str | ZarrV3NamedConfigJSON +ZarrV3MetadataFieldJSON: TypeAlias = str | ZarrV3NamedConfigJSON """The JSON shape of any v3 metadata extension-point entry: either a bare short-hand name string or a `{name, configuration, must_understand}` envelope. @@ -30,18 +36,84 @@ """ +# A member holding a metadata field is annotated with the alias of its +# kind, which a scope reads it as. Each alias is the JSON a field is, so to +# a type checker, and to `check`, it is `ZarrV3MetadataFieldJSON`. + +DataTypeField = TypeAliasType("DataTypeField", ZarrV3MetadataFieldJSON) +"""A member holding a data type: a document's `data_type`, or a struct field's; read in the scope what holds it is read in.""" +ChunkGridField = TypeAliasType("ChunkGridField", ZarrV3MetadataFieldJSON) +"""A member holding a chunk grid: a document's `chunk_grid`.""" +ChunkKeyEncodingField = TypeAliasType("ChunkKeyEncodingField", ZarrV3MetadataFieldJSON) +"""A member holding a chunk key encoding: a document's `chunk_key_encoding`.""" +CodecField = TypeAliasType("CodecField", ZarrV3MetadataFieldJSON) +"""A member holding a codec: a document's `codecs` is `tuple[CodecField, ...]`, and so is a shard's.""" +StaticCodecField = TypeAliasType("StaticCodecField", ZarrV3MetadataFieldJSON) +"""A member holding a codec of static size: a shard's `index_codecs` is one, +since a reader finds the index by a size it knows before reading it. +""" +StorageTransformerField = TypeAliasType("StorageTransformerField", ZarrV3MetadataFieldJSON) +"""A member holding a storage transformer: a document's `storage_transformers` is `tuple[StorageTransformerField, ...]`.""" + + def validate_metadata_field_v3( value: object, *, allow_must_understand_false: bool = True ) -> tuple[ValidationProblem, ...]: """Return every reason `value` is not a v3 metadata field. - A metadata field is a bare name string or a mapping containing `name` and - optional `configuration` and `must_understand` members: an envelope, as - `envelope_problems` judges it, around a configuration whose members are - JSON. + A metadata field is a bare name, or an envelope around a configuration + whose members are JSON: an object of a `name` as the spec names an + extension, a `configuration` that is an object of string keys, a + boolean `must_understand`, and nothing else. """ envelope = envelope_problems(value, allow_must_understand_false=allow_must_understand_false) - return (*envelope, *_configuration_json_problems(value)) + return with_input((*envelope, *_configuration_json_problems(value)), value) + + +ENVELOPE_KEYS: frozenset[str] = frozenset({"name", "configuration", "must_understand"}) +"""The members a metadata field's envelope declares: a key beside them is an unknown key.""" + + +_EXTENSION_NAME: Final = re.compile(r"[a-z][a-z0-9_.-]+") +"""A registered extension name: the spec's `^[a-z][a-z0-9-_.]+$`, its hyphen last so no engine reads it as a range.""" + +_URI: Final = re.compile(r"[A-Za-z][A-Za-z0-9+.-]*:[A-Za-z0-9._~:/?#\[\]@!$&'()*+,;=%-]+") +"""A URI, as far as a name is judged: RFC 3986's scheme, its colon, and one or more of the characters a URI is written in -- its unreserved and reserved sets, and the `%` of percent-encoding (https://www.rfc-editor.org/rfc/rfc3986#section-2) -- the names earlier versions of the spec required. + +The characters are listed rather than written `\\S`, which Python's `re` +and ECMA-262, the dialect a JSON Schema `pattern` is read in, disagree +on: `\\x1c`-`\\x1f` and `\\x85` are whitespace to one and `\\ufeff` to the +other, so a name this reader refused, a validator of the schema would +accept, or the other way round. +""" + +EXTENSION_NAME_SCHEMA_PATTERN: Final = rf"^({_EXTENSION_NAME.pattern}|{_URI.pattern})(?![\s\S])" +r"""What `well_named` accepts, as a JSON Schema `pattern`: the two patterns it matches whole, `(?![\s\S])` where `$` would take a final newline, as `_RAW_BYTES_SCHEMA_PATTERN` has it.""" + + +def well_named(name: str) -> bool: + """Whether `name` is an extension name as the spec names one. + + A registered name "MUST start with one lower case letter a-z and then + be followed by only lower case letters a-z, numerals 0-9, underscores, + dots and dashes", regex `^[a-z][a-z0-9-_.]+$` + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L1603-L1606); + a URI, which earlier versions of the spec required, is "still + permitted" + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L1614-L1618). + """ + return _EXTENSION_NAME.fullmatch(name) is not None or _URI.fullmatch(name) is not None + + +def name_problem(name: str, at: tuple[str | int, ...]) -> ValidationProblem | None: + """The problem `name`, at `at`, is when the spec gives no extension such a name; None when it does, as `well_named` says.""" + if well_named(name): + return None + message = ( + "expected an extension name -- lower-case letters, digits, '-', '_' and '.', " + f"starting with a letter -- or a URI, got {shown(name)}" + ) + return ValidationProblem(at, message, "invalid_value") def envelope_problems( @@ -49,42 +121,47 @@ def envelope_problems( ) -> tuple[ValidationProblem, ...]: """Every reason `value` is not a v3 metadata field's envelope, what its configuration holds left unjudged. - The envelope is what sits around the configuration: a string `name`, a - `configuration` that is an object of string keys, a boolean - `must_understand`, and nothing else. `resolve` asks this of a field it - has refined to JSON already, so a configuration is walked once. + The envelope is what sits around the configuration: a `name` as the + spec names an extension, a `configuration` that is an object of string + keys, a boolean `must_understand`, and nothing else. `resolve` asks + this of a field it has refined to JSON already, so a configuration is + walked once. """ if isinstance(value, str): - return () - if not isinstance(value, Mapping): + bad = name_problem(value, ()) + return () if bad is None else (bad,) + if not is_object(value): return ( ValidationProblem( (), - "expected a metadata field (string or extension object)", + f"expected a metadata field (string or extension object), got {shown(value)}", "invalid_type", ), ) - field = cast("Mapping[object, object]", value) + field = value problems: list[ValidationProblem] = [] - allowed_keys = frozenset({"name", "configuration", "must_understand"}) for key in field: if not isinstance(key, str): problems.append( - ValidationProblem((), f"non-string metadata field key {key!r}", "invalid_type") - ) - elif key not in allowed_keys: - problems.append( - ValidationProblem((key,), "unexpected metadata field member", "invalid_value") + ValidationProblem( + (), f"non-string metadata field key {shown_key(key)}", "invalid_type" + ) ) - if not isinstance(field.get("name"), str): + elif key not in ENVELOPE_KEYS: + problems.append(ValidationProblem((key,), f"unexpected key {key!r}", "unknown_key")) + if "name" not in field: + problems.append(ValidationProblem(("name",), "missing required key", "missing_key")) + elif not isinstance(field["name"], str): problems.append(ValidationProblem(("name",), "expected a string name", "invalid_type")) + elif (bad := name_problem(field["name"], ("name",))) is not None: + problems.append(bad) if "configuration" in field: configuration = field["configuration"] - if not isinstance(configuration, Mapping): + if not is_object(configuration): problems.append( - ValidationProblem(("configuration",), "expected a mapping", "invalid_type") + ValidationProblem(("configuration",), "expected an object", "invalid_type") ) - elif not all(isinstance(k, str) for k in cast("Mapping[object, object]", configuration)): + elif not all(isinstance(k, str) for k in configuration): problems.append( ValidationProblem(("configuration",), "expected string keys", "invalid_type") ) @@ -107,44 +184,53 @@ def envelope_problems( def _configuration_json_problems(value: object) -> tuple[ValidationProblem, ...]: """Each member of `value`'s configuration that is not JSON, located; nothing where no object of string keys is there to walk.""" - if not isinstance(value, Mapping): + if not is_object(value): return () - configuration = cast("Mapping[object, object]", value).get("configuration") - if not isinstance(configuration, Mapping): + configuration = value.get("configuration") + if not is_object(configuration): return () - members = cast("Mapping[object, object]", configuration) + members = configuration if not all(isinstance(key, str) for key in members): return () return tuple( found - for key, item in cast("Mapping[str, object]", members).items() - for found in prefixed("configuration", prefixed(key, validate_json(item))) + for key, item in members.items() + if isinstance(key, str) + for found in validate_json(item, ("configuration", key)) ) def is_metadata_field_v3(value: object) -> TypeGuard[ZarrV3MetadataFieldJSON]: - """Whether `value` is a v3 metadata field: a bare name or a named config.""" + """Whether `value` is a v3 metadata field: a bare name as the spec names an extension, or a named config.""" if isinstance(value, str): - return True - if not isinstance(value, dict): + return well_named(value) + if not is_object(value) or not isinstance(value, dict): return False - field = cast("dict[object, object]", value) - return is_canonical_json(field) and not validate_metadata_field_v3(field) + return is_canonical_json(value) and not validate_metadata_field_v3(value) def parse_metadata_field_v3(value: object) -> ZarrV3MetadataFieldJSON: """Return `value` narrowed to `ZarrV3MetadataFieldJSON`, or raise `MetadataValidationError`.""" - normalized = arrays_to_tuples(value) - problems = validate_metadata_field_v3(normalized) + problems = validate_metadata_field_v3(value) if len(problems) != 0: raise MetadataValidationError(problems) - return cast(ZarrV3MetadataFieldJSON, normalized) + return cast(ZarrV3MetadataFieldJSON, arrays_to_tuples(value)) __all__ = [ + "ENVELOPE_KEYS", + "EXTENSION_NAME_SCHEMA_PATTERN", + "ChunkGridField", + "ChunkKeyEncodingField", + "CodecField", + "DataTypeField", + "StaticCodecField", + "StorageTransformerField", "ZarrV3MetadataFieldJSON", "envelope_problems", "is_metadata_field_v3", + "name_problem", "parse_metadata_field_v3", "validate_metadata_field_v3", + "well_named", ] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/_definition.py b/packages/zarr-metadata/src/zarr_metadata/v3/_definition.py index c18a0bd260..4457eb8695 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/_definition.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/_definition.py @@ -13,7 +13,7 @@ needs none. 2. The rules need the definition: plain functions over the checked TypedDict, for everything finer than a type -- a bound, members read - together. `Definition.judge` is the check and then the rules, for a + together. `Definition.read_configuration` is the check and then the rules, for a caller holding one configuration. 3. The reading needs a scope. `resolve` relates the field's name to a definition through a `Context`, judges the configuration, and reads @@ -34,12 +34,15 @@ import re from collections.abc import Callable, Iterable, Iterator, Mapping from dataclasses import dataclass +from types import MappingProxyType from typing import ( TYPE_CHECKING, Any, + ClassVar, Final, Generic, Literal, + Protocol, TypeAlias, TypeGuard, cast, @@ -51,18 +54,44 @@ from typing_extensions import TypeAliasType, TypedDict, TypeVar, is_typeddict from zarr_metadata._common import JSONValue, ZarrV3NamedConfigJSON -from zarr_metadata._json import ValidationProblem, refine_json +from zarr_metadata._json import ( + ValidationProblem, + copied, + frozen, + is_object, + is_tuple, + json_text, + refine_json, + shown, + with_input, +) +from zarr_metadata._sentinel import UNSET from zarr_metadata._typed_json import ( + JSONSchema, Loc, Parsed, Parser, - no_leaf, + SchemaLeaf, + Schemas, + is_alias, parser, problem, typeddict_keys, unread_in, ) -from zarr_metadata.v3._common import ZarrV3MetadataFieldJSON, envelope_problems +from zarr_metadata.v3._common import ( + ENVELOPE_KEYS, + EXTENSION_NAME_SCHEMA_PATTERN, + ChunkGridField, + ChunkKeyEncodingField, + CodecField, + DataTypeField, + StaticCodecField, + StorageTransformerField, + envelope_problems, + name_problem, + well_named, +) if TYPE_CHECKING: from collections.abc import Sequence @@ -80,6 +109,7 @@ T = TypeVar("T") D = TypeVar("D", bound="Definition[Any]") +F = TypeVar("F", bound="WithFillValue[Any]") Problems: TypeAlias = tuple[ValidationProblem, ...] @@ -94,6 +124,11 @@ def unchanged(configuration: T) -> T: return configuration +def fill_value_as_written(configuration: object, nested: object, value: T) -> T: + """The canonical spelling of a fill value of a data type that spells each of its values one way: the fill value as written.""" + return value + + def unknown_lengths(configuration: object, nested: object, shape: tuple[int, ...]) -> Lengths: """The chunk lengths of a grid that says nothing of them: unknown, along each axis of the array.""" return (None,) * len(shape) @@ -139,7 +174,11 @@ def unknown_storage(*_: object) -> None: class EmptyConfiguration(TypedDict, closed=True): - """The configuration of a definition with nothing to configure, written as its bare name.""" + """The configuration of a definition with nothing to configure: its field is written with its name alone.""" + + +_FIELD_KINDS: Final[dict[TypeAliasType, type[Definition[Any]]]] = {} +"""Each field alias, and the kind a member annotated with it is read as; filed by each kind as its class is built.""" @dataclass(frozen=True, kw_only=True, slots=True) @@ -156,9 +195,15 @@ class Definition(Generic[C]): type; with `closed=False`, anything. `rules` yields what the spec disallows in a configuration of that type, as it finds each; it is handed only a configuration that has passed the check, holding what - the TypedDict admits and nothing else, and the fields it holds as the - scope read them -- a struct's field types -- which is nothing when no - scope read it. `judge` is the two, for a caller holding JSON. + the TypedDict admits and nothing else, each member within the bounds + its type carries, and the fields it holds as the scope read them -- a + struct's field types -- which is nothing when no scope read it. + `read_configuration` is the two, for a caller holding JSON. + + Each function is handed the configuration as a read-only view, no + `dict`: `copy.deepcopy` and `json.dumps` refuse it, and a function that + folds a spelling builds a new mapping, `{**configuration}` without the + member, rather than editing what it was handed. `canonical` is where two spellings of the configuration that mean the same thing are made one. @@ -169,6 +214,19 @@ class Definition(Generic[C]): and a `name` or rules that are not what they say. """ + _is_kind: ClassVar[bool] = False + """Whether this class is a kind, what a scope files definitions by: set by `class Kind(Definition, kind=True, format=3)`. + + Read from a class's own namespace, never inherited: a subclass of a + kind is a definition of that kind, and declares no `kind`. + """ + _format: ClassVar[Literal[2, 3] | None] = None + """The Zarr format whose documents hold fields of this kind, declared with the kind and inherited by its definitions: what a scope reads documents of.""" + label: ClassVar[str] = "definition" + """The kind as a message names it: "codec".""" + field_aliases: ClassVar[tuple[TypeAliasType, ...]] = () + """The field aliases a configuration member holding a field of this kind is annotated with: `CodecField` and `StaticCodecField` for a codec.""" + name: str """The name the metadata carries, which a scope files the definition under.""" configuration: type[C] @@ -182,6 +240,86 @@ class Definition(Generic[C]): canonical form by `canonicalize`, which knows where each one sits. """ + @classmethod + def name_problem(cls, name: str, at: Loc) -> ValidationProblem | None: + """The problem `name`, at `at`, is when no document of the format writes it for a field of this kind; None when one may: for Zarr v3, when the spec gives an extension such a name.""" + return name_problem(name, at) + + @classmethod + def well_named(cls, name: str) -> bool: + """Whether a document of the format may write `name` for a field of this kind, as `name_problem` says.""" + return cls.name_problem(name, ()) is None + + @classmethod + def spelled(cls, name: str) -> tuple[str | None, dict[str, JSONValue] | None]: + """How a name a document writes reads: the name its definition is filed under, and the configuration the name carries. + + A name is filed as itself and carries nothing, `(name, None)`. A + kind whose names carry configuration says otherwise: a v3 data + type `r16` is filed under `r*` with `{"bits": 16}`. A name no + document writes, which only files a definition, is `(None, None)`. + """ + return name, None + + def carrying_name(self, configuration: Mapping[str, JSONValue]) -> str | None: + """The name that carries `configuration` for this definition, the inverse of `spelled`: `r16` for `r*` with `{"bits": 16}`; None when its names carry nothing.""" + return None + + @classmethod + def named_configuration( + cls, value: object + ) -> tuple[str | None, Mapping[str, object] | None, Problems]: + """`value`, a field as a document of the format writes it, split into `(name, configuration, problems)`, as the module's `named_configuration` splits a v3 field.""" + return named_configuration(value) + + @classmethod + def envelope_problems(cls, value: object) -> Problems: + """Every reason `value` is not a field's envelope as the format writes one for this kind, what the configuration holds left unjudged.""" + return envelope_problems(value, allow_must_understand_false=False) + + @classmethod + def envelope_json(cls, name: str, configuration: Mapping[str, JSONValue]) -> JSONValue: + """A field of `name` and `configuration` as a document of the format writes it, in the fewest words every reader takes: for v3, an object.""" + if len(configuration) != 0: + return {"name": name, "configuration": configuration} + return {"name": name} + + @classmethod + def configuration_loc(cls, loc: Loc) -> Loc: + """Where the configuration of a field at `loc` sits: under `configuration` for v3; at the field for a format that writes the parameters beside the name.""" + return (*loc, "configuration") + + @classmethod + def name_loc(cls, loc: Loc) -> Loc: + """Where the name of a field at `loc`, written as an object, sits: under `name` for v3.""" + return (*loc, "name") + + def __init_subclass__( + cls, *, kind: bool = False, format: Literal[2, 3] | None = None, **kwargs: object + ) -> None: + """Files a subclass: `kind=True` declares a kind, of the Zarr `format` whose documents hold its fields, as `typing.Protocol` and SQLAlchemy's `__abstract__` mark a class and not its subclasses.""" + # Named, not `super()`: a dataclass with slots is rebuilt, and the + # cell a bare `super()` reads names the class that was thrown away. + super(Definition, cls).__init_subclass__(**kwargs) + # A dataclass with slots is built twice, the second time without + # the class keywords but with the first class's namespace, and the + # class built last is the one a document is read with: the mark is + # kept in the namespace, and the aliases are filed last. + if kind: + if format not in (2, 3): + msg = ( + f"{cls.__name__}: a kind declares the Zarr format its fields belong to, " + f"format=2 or format=3, got {format!r}" + ) + raise TypeError(msg) + cls._is_kind = True + cls._format = format + elif format is not None: + msg = f"{cls.__name__}: only a kind declares a format; a definition of a kind has its kind's" + raise TypeError(msg) + for alias in cls.__dict__.get("field_aliases", ()): + _FIELD_KINDS[alias] = cls + def __post_init__(self) -> None: refusal = _malformed(self) or self._refusal() if refusal is not None: @@ -192,6 +330,11 @@ def __post_init__(self) -> None: msg = f"{self.name!r}: {error}" raise TypeError(msg) from error + def __repr__(self) -> str: + # Short, as a reading that holds definitions shows them: in full, a + # definition's repr is each function it holds, at its address. + return f"{type(self).__name__}(name={self.name!r})" + def _refusal(self) -> str | None: """What is wrong with the members a kind adds; None when nothing is, or it adds none.""" return None @@ -201,31 +344,35 @@ def requires_configuration(self) -> bool: """Whether a document must write a configuration: whether the TypedDict has a required key.""" return len(typeddict_keys(self.configuration).required) != 0 - def check(self, value: object, loc: Loc = ()) -> tuple[C | None, Problems]: + def _check_configuration(self, value: object, loc: Loc = ()) -> tuple[C | None, Problems]: """`value` type-checked as this definition's configuration, each nested field's envelope judged. `zarr_metadata.typed_json.check` is the type check alone; this also judges the envelope of each metadata field a member holds. """ - return _configuration_checked(value, self.configuration, loc) + configuration, problems = _configuration_checked(value, self.configuration, loc) + return configuration, with_input(problems, value, loc) - def judge(self, value: object, loc: Loc = ()) -> tuple[C | None, Problems]: + def read_configuration(self, value: object, loc: Loc = ()) -> tuple[C | None, Problems]: """`value` type-checked, then judged by the rules: the configuration if it holds, and every problem. - The rules are asked only of a configuration that type-checked and - whose nested fields are well formed, holding what its TypedDict - admits and nothing else, so a caller holding JSON never reaches a - rule with a member of the wrong type, or one the type says cannot - be there. No scope reads the fields it holds, so the rules see + The rules are asked only of a configuration that type-checked, + its bounds kept, and whose nested fields are well formed, holding + what its TypedDict admits and nothing else, so a caller holding + JSON never reaches a rule with a member of the wrong type, out of + its bounds, or one the type says cannot be there. No scope reads the fields it holds, so the rules see none of them read, and a rule about one -- a struct's field of a type whose values vary in size -- finds nothing to judge: `resolve` reads the field in a scope, and asks every rule. """ - configuration, problems = self.check(value, loc) + configuration, problems = self._check_configuration(value, loc) if configuration is None: return None, problems - refused = ruled(self, lambda: self.rules(configuration, _nothing_nested()), loc) - return (configuration if len(refused) == 0 else None), (*problems, *refused) + refused = ruled(self, lambda: self.rules(read_only(configuration), _nothing_nested()), loc) + return (configuration if len(refused) == 0 else None), ( + *problems, + *with_input(refused, value, loc), + ) def _malformed(definition: Definition[Any]) -> str | None: @@ -237,6 +384,14 @@ def _malformed(definition: Definition[Any]) -> str | None: value = getattr(definition, member) if not callable(value): return f"{name!r}: {member} is a function, got {value!r}" + kind = type(definition) + filed, _ = kind.spelled(name) + bad = None if filed is None else kind.name_problem(name, ()) + if bad is not None: + return ( + f"{name!r}: {bad.message}, so no document names it, and nothing would ever read " + "with this definition" + ) return None @@ -260,14 +415,16 @@ def _function_members(kind: type[Definition[Any]]) -> tuple[str, ...]: (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/data-types/index.rst#L46-L47). """ -RAW_BYTES_NAME_PATTERN: Final = re.compile(r"r([0-9]+)") -"""A name that writes raw bits: `r` and a size in bits, matched whole, whether or not the size is allowed. +RAW_BYTES_NAME_PATTERN: Final = re.compile(r"r([0-9]{1,100})") +"""A name that writes raw bits: `r` and a size in bits of up to a hundred digits, matched whole, whether or not the size is allowed. ASCII digits only: `\\d` would also match every other Unicode decimal, so `r\uff11\uff16` would be read as sixteen bits, and a third-party name spelled that way would be taken for raw bits. A size the spec does not allow -- `r0`, `r12` -- is still raw bits, so it is reported as a bad -size rather than passed as an unknown extension. +size rather than passed as an unknown extension. A hundred digits is +more than any size: `r` and more digits than that is a name, as the spec +allows one, which nothing in scope claims. """ @@ -281,36 +438,62 @@ def spelled( itself is filed under nothing, `(None, None)`, since no document writes it. """ - if kind is DataTypeDefinition: - if name == RAW_BYTES_NAME: - return None, None - written = RAW_BYTES_NAME_PATTERN.fullmatch(name) - if written is not None: - return RAW_BYTES_NAME, {"bits": int(written.group(1))} - return name, None + return kind.spelled(name) + + +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class WithFillValue(Definition[C]): + """A definition of a type whose arrays take a fill value: the JSON shape of one, what the rules disallow in one, and its canonical spelling. + + Not a kind: each format's data type kind derives from it, so + `fill_value_problems` judges a fill value of either. + `fill_value` is the JSON shape of a fill value, as an annotation the + checker reads as it reads a configuration's members, its range among + it. `fill_value_rules` is what the spec disallows in a fill value of + that shape that the type cannot say; it is handed the configuration, + the fields it holds as the scope read them, and the typed fill value. + `fill_value_canonical` spells a fill value that has no problem in the + one spelling its value has, so two fill values are one value of the + type exactly when their canonical spellings are written alike. + """ -def _carrying_name( - definition: Definition[Any], configuration: Mapping[str, JSONValue] -) -> str | None: - """The name that carries `configuration` for `definition`: `r16` for `r*` with `{"bits": 16}`; None when its name carries nothing.""" - if not isinstance(definition, DataTypeDefinition) or definition.name != RAW_BYTES_NAME: + fill_value: object = JSONValue + """The JSON shape of a fill value, as an annotation: `Int8FillValue`.""" + fill_value_rules: Callable[[C, Nested, Any], Iterable[ValidationProblem]] = no_rules + """What the spec disallows in a fill value of that shape, located in it.""" + fill_value_canonical: Callable[[C, Nested, Any], JSONValue] = fill_value_as_written + """A fill value that has no problem, in the one spelling its value has.""" + + def _refusal(self) -> str | None: + try: + _fill_value_parser(self.fill_value) + except TypeError as error: + return f"{self.name!r}: fill_value: {error}" return None - return f"r{configuration['bits']}" -@dataclass(frozen=True, kw_only=True, slots=True) -class DataTypeDefinition(Definition[C]): +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class DataTypeDefinition(WithFillValue[C], kind=True, format=3): """A data type, and the fill value an array of it takes. `fill_value` is the JSON shape of a fill value -- `Int8FillValue`, an - annotation the checker reads as it reads a configuration's members -- - and `fill_value_rules` is what the spec disallows in a fill value of - that shape: an integer out of range, a hex string of another width. + annotation the checker reads as it reads a configuration's members, + its range among it -- and `fill_value_rules` is what the spec + disallows in a fill value of that shape that the type cannot say: a + hex string of another width. The rules are handed the configuration, the fields it holds as the scope read them (a struct's field types), and the typed fill value. A data type that says nothing of its fill value takes any JSON. + `fill_value_canonical` spells a fill value that has no problem -- + well typed, and allowed by the rules -- in the one spelling its value + has, so two fill values are one value of the type exactly when their + canonical spellings are written alike: `"NaN"` and `"0x7fc00000"` are + one `float32`, and `0.0` and `-0.0` two. It is handed what the rules + are handed. A data type that says nothing of it spells each of its + values one way: as written. + `storage` says how its values are stored -- in single bytes, in several bytes at a time, or each in as many as it needs -- which is what the `bytes` codec asks of the data type it is handed: an @@ -323,10 +506,31 @@ class DataTypeDefinition(Definition[C]): this definition. """ - fill_value: object = JSONValue - """The JSON shape of a fill value, as an annotation: `Int8FillValue`.""" - fill_value_rules: Callable[[C, Nested, Any], Iterable[ValidationProblem]] = no_rules - """What the spec disallows in a fill value of that shape, located in it.""" + label: ClassVar[str] = "data type" + field_aliases: ClassVar[tuple[TypeAliasType, ...]] = (DataTypeField,) + + @classmethod + def spelled(cls, name: str) -> tuple[str | None, dict[str, JSONValue] | None]: + if name == RAW_BYTES_NAME: + return None, None + written = RAW_BYTES_NAME_PATTERN.fullmatch(name) + if written is not None: + return RAW_BYTES_NAME, {"bits": int(written.group(1))} + return name, None + + def carrying_name(self, configuration: Mapping[str, JSONValue]) -> str | None: + if self.name != RAW_BYTES_NAME: + return None + return f"r{configuration['bits']}" + + @classmethod + def envelope_json(cls, name: str, configuration: Mapping[str, JSONValue]) -> JSONValue: + # A data type with nothing to configure is its bare name, as core + # data types have been written since Zarr v3.0. + if len(configuration) != 0: + return {"name": name, "configuration": configuration} + return name + storage: Callable[[C, Nested], StorageClass | None] = unknown_storage """How its values are stored, given the configuration and the fields it holds; None when unknown.""" @@ -336,21 +540,18 @@ def _refusal(self) -> str | None: f"{self.name!r} is how a document writes raw bits of one size, which read as " f"{RAW_BYTES_NAME!r}; to read raw bits your own way, define {RAW_BYTES_NAME!r}" ) - try: - _fill_value_parser(self.fill_value) - except TypeError as error: - return f"{self.name!r}: fill_value: {error}" - return None + # Named, not a bare `super()`: a dataclass with slots is rebuilt. + return super(DataTypeDefinition, self)._refusal() @functools.cache def _fill_value_parser(annotation: object) -> Parser: - """The checker for a fill value's JSON shape, compiled once; `TypeError` naming what no checker reads.""" - return parser(annotation, no_leaf) + """The checker for a fill value's JSON shape, compiled once; `TypeError` naming what no checker reads, or a metadata field in it.""" + return parser(annotation, _no_field) -@dataclass(frozen=True, kw_only=True, slots=True) -class ChunkGridDefinition(Definition[C]): +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class ChunkGridDefinition(Definition[C], kind=True, format=3): """A chunk grid, and the arrays it fits. `shape_rules` is what the spec disallows in a grid of this @@ -368,16 +569,22 @@ class ChunkGridDefinition(Definition[C]): every axis unknown. """ + label: ClassVar[str] = "chunk grid" + field_aliases: ClassVar[tuple[TypeAliasType, ...]] = (ChunkGridField,) + shape_rules: Callable[[C, Nested, tuple[int, ...]], Iterable[ValidationProblem]] = no_rules """What the spec disallows in this grid over an array of a shape, located in the configuration.""" chunk_lengths: Callable[[C, Nested, tuple[int, ...]], Lengths] = unknown_lengths """The lengths its chunks take along each axis of an array of a shape it fits, None where unknown.""" -@dataclass(frozen=True, kw_only=True, slots=True) -class ChunkKeyEncodingDefinition(Definition[C]): +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class ChunkKeyEncodingDefinition(Definition[C], kind=True, format=3): """A chunk key encoding.""" + label: ClassVar[str] = "chunk key encoding" + field_aliases: ClassVar[tuple[TypeAliasType, ...]] = (ChunkKeyEncodingField,) + CodecKind = Literal["array_array", "array_bytes", "bytes_bytes"] """What a codec does to what it is handed: the three positions a pipeline orders.""" @@ -397,8 +604,8 @@ class ChunkKeyEncodingDefinition(Definition[C]): """The functions no codec of a kind is asked: a bytes -> bytes codec is handed bytes, and only an array -> array codec hands on a chunk.""" -@dataclass(frozen=True, kw_only=True, slots=True) -class CodecDefinition(Definition[C]): +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class CodecDefinition(Definition[C], kind=True, format=3): """A codec: what it does to what it is handed, and whether the size of what it gives out is static. A codec handed an array -- array -> array, array -> bytes -- says what @@ -429,6 +636,9 @@ class CodecDefinition(Definition[C]): codec, which is handed bytes -- is refused. """ + label: ClassVar[str] = "codec" + field_aliases: ClassVar[tuple[TypeAliasType, ...]] = (CodecField, StaticCodecField) + kind: CodecKind size: CodecSize chunk_rules: Callable[[C, Nested, Chunk], Iterable[ValidationProblem]] = no_rules @@ -452,10 +662,13 @@ def _refusal(self) -> str | None: return None -@dataclass(frozen=True, kw_only=True, slots=True) -class StorageTransformerDefinition(Definition[C]): +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class StorageTransformerDefinition(Definition[C], kind=True, format=3): """A storage transformer.""" + label: ClassVar[str] = "storage transformer" + field_aliases: ClassVar[tuple[TypeAliasType, ...]] = (StorageTransformerField,) + KINDS: Final[tuple[type[Definition[Any]], ...]] = ( DataTypeDefinition, @@ -468,55 +681,64 @@ class StorageTransformerDefinition(Definition[C]): def kind_of(definition: Definition[Any]) -> type[Definition[Any]] | None: - """The kind `definition` is; None for a definition of no kind, which no scope files.""" - return next((kind for kind in KINDS if isinstance(definition, kind)), None) + """The kind `definition` is: the nearest class in its MRO declared with `kind=True`; None for a definition of no kind, which no scope files.""" + bases: tuple[type[object], ...] = type(definition).__mro__ + return next((base for base in bases if _declares_kind(base)), None) + + +def format_of(kind: type[Definition[Any]]) -> Literal[2, 3]: + """The Zarr format whose documents hold fields of `kind`, as the kind declared it; `TypeError` for a class that is no kind.""" + found = as_kind(kind)._format # pyright: ignore[reportPrivateUsage] + if found is None: # a kind cannot be declared without one + msg = f"{kind!r} declares no format" + raise TypeError(msg) + return found + + +def _declares_kind(cls: type[object]) -> TypeGuard[type[Definition[Any]]]: + """Whether `cls` itself was declared `kind=True`: the mark is read from its own namespace, which a subclass does not share.""" + return vars(cls).get("_is_kind") is True and issubclass(cls, Definition) def as_kind(kind: object) -> type[Definition[Any]]: """The kind of metadata `kind` names, type arguments dropped; `TypeError` if it names none. - A scope files definitions by kind, so a field is read as one of - `KINDS` -- `CodecDefinition`, or `CodecDefinition[Any]` -- and never - as the base `Definition` or a class of the caller's own, under which - nothing is filed: a field read as one would go unjudged. + A scope files definitions by kind, so a field is read as a kind -- + `CodecDefinition`, or `CodecDefinition[Any]` -- and never as the base + `Definition` or a class that declares no kind, under which nothing is + filed: a field read as one would go unjudged. """ origin = get_origin(kind) or kind - found = next((known for known in KINDS if origin is known), None) - if found is None: - names = ", ".join(known.__name__ for known in KINDS) - msg = f"{kind!r} is not a kind of metadata; read a field as one of {names}" - raise TypeError(msg) - return found - + if isinstance(origin, type) and _declares_kind(origin): + return origin + names = ", ".join(known.__name__ for known in KINDS) + msg = ( + f"{kind!r} is not a kind of metadata; read a field as one of {names}, or as a " + "subclass of Definition declared with kind=True and its format" + ) + raise TypeError(msg) -DataTypeField = TypeAliasType("DataTypeField", ZarrV3MetadataFieldJSON) -"""A configuration member holding a data type, read in the scope the member's field is read in.""" -ChunkGridField = TypeAliasType("ChunkGridField", ZarrV3MetadataFieldJSON) -"""A configuration member holding a chunk grid.""" -ChunkKeyEncodingField = TypeAliasType("ChunkKeyEncodingField", ZarrV3MetadataFieldJSON) -"""A configuration member holding a chunk key encoding.""" -CodecField = TypeAliasType("CodecField", ZarrV3MetadataFieldJSON) -"""A configuration member holding a codec: a shard's `codecs` is `tuple[CodecField, ...]`.""" -StaticCodecField = TypeAliasType("StaticCodecField", ZarrV3MetadataFieldJSON) -"""A configuration member holding a codec of static size: a shard's `index_codecs` is one, -since a reader finds the index by a size it knows before reading it. -""" -StorageTransformerField = TypeAliasType("StorageTransformerField", ZarrV3MetadataFieldJSON) -"""A configuration member holding a storage transformer.""" - -_FIELD_KINDS: Final[Mapping[object, type[Definition[Any]]]] = { - DataTypeField: DataTypeDefinition, - ChunkGridField: ChunkGridDefinition, - ChunkKeyEncodingField: ChunkKeyEncodingDefinition, - CodecField: CodecDefinition, - StaticCodecField: CodecDefinition, - StorageTransformerField: StorageTransformerDefinition, -} -_STATIC_SIZE: Final[frozenset[object]] = frozenset({StaticCodecField}) +_STATIC_SIZE: Final[frozenset[TypeAliasType]] = frozenset({StaticCodecField}) """The field aliases whose codec must be of static size.""" +def field_kind(annotation: object) -> type[Definition[Any]] | None: + """The kind of metadata field a member annotated `annotation` holds -- `CodecDefinition` for `CodecField` -- or None when it holds none.""" + if not is_alias(annotation): + return None + return _FIELD_KINDS.get(annotation) + + +class Scope(Protocol): + """What reading a field asks of a scope: its format, and which definition of a kind claims a name. A `Context` is one; so is anything else that answers the two.""" + + @property + def format(self) -> Literal[2, 3] | None: ... + + def claimant(self, kind: type[D], name: str) -> D | None: ... + + @dataclass(frozen=True, slots=True) class _NestedField: """A metadata field the check met inside a configuration: where it sits, its kind, its JSON. @@ -541,18 +763,14 @@ def _field(annotation: object) -> Parser | None: judged, and the name related to a definition, by whoever reads the field -- `check` without a scope, `resolve` in one. """ - try: - kind = _FIELD_KINDS.get(annotation) - except TypeError: # an unhashable annotation is no field alias - return None + kind = field_kind(annotation) if kind is None: return None static = annotation in _STATIC_SIZE def parse(value: object, loc: Loc) -> Parsed: - if not isinstance(value, (str, Mapping)): - return value, problem(loc, f"expected a metadata field, got {value!r}") - # Refined JSON, which the checker only knows as `object`. + # Refined JSON, which the checker only knows as `object`; what is + # not a field at all, the kind's envelope says. return _NestedField(loc, kind, cast("JSONValue", value), static), () return parse @@ -571,7 +789,7 @@ def _vetting(annotation: object) -> Parser | None: msg = ( "a member typed ZarrV3MetadataFieldJSON checks as plain JSON, and is never read as " "a field; annotate it with the field alias of its kind: " - + ", ".join(cast("TypeAliasType", alias).__name__ for alias in _FIELD_KINDS) + + ", ".join(alias.__name__ for alias in _FIELD_KINDS) ) raise TypeError(msg) if ( @@ -589,6 +807,15 @@ def _vetting(annotation: object) -> Parser | None: return _field(annotation) +def _no_field(annotation: object) -> Parser | None: + """A leaf refusing a field alias: a fill value is a value of its data type, and holds no metadata field.""" + if is_alias(annotation) and field_kind(annotation) is not None: + name = annotation.__name__ + msg = f"{name} holds a metadata field, and a fill value is a value of its data type" + raise TypeError(msg) + return None + + @functools.cache def _vet(configuration: type) -> None: """Refuse a configuration no definition could read with, saying what is wrong with it.""" @@ -604,6 +831,127 @@ def _vet(configuration: type) -> None: raise TypeError(msg) +def kind_field(kind: type[Definition[Any]]) -> object: + """The field alias a member holding any field of `kind` is annotated with: the first the kind declares that does not narrow the field, `CodecField` rather than `StaticCodecField`.""" + return next(alias for alias in kind.field_aliases if alias not in _STATIC_SIZE) + + +_RAW_BYTES_SCHEMA_PATTERN: Final = f"^{RAW_BYTES_NAME_PATTERN.pattern}(?![\\s\\S])" +"""`RAW_BYTES_NAME_PATTERN`, matched whole, as a JSON Schema writes a pattern. + +Held to the end of the name by a lookahead for no character at all: a +`$` there would also match before a final newline in a validator that +matches patterns as Python does, so `"r16\\n"`, which names nothing, +would read as raw bits. +""" + + +def field_json_schema(kind: type[Definition[Any]], context: Context) -> JSONSchema: + """The JSON Schema of one metadata field read as `kind` in `context`: what `resolve` reads, but for the rules. + + A field one of the definitions in scope reads -- its name, its + configuration as the TypedDict says, a `must_understand` of `true` + if any, and its bare name when it needs no configuration -- or a + name none of them claims, with any configuration: what keeps the + format open. A field a configuration holds is written the same way, + in the same scope, and a member taking codecs of static size only + takes those. JSON Schema draft 2020-12, as `json_schema` writes one; + the fields it holds, and the configuration of each definition, are in + `$defs`, under the name of the field alias or TypedDict. What only a + rule says -- a blosc `typesize` against its `shuffle` -- is not in it, + so a field it accepts may still have a problem. + """ + asked = as_kind(kind) + if not any(issubclass(asked, known) for known in KINDS): + msg = ( + f"{asked.__name__} is a kind of another format; the JSON Schema writer writes " + "Zarr v3 fields only" + ) + raise TypeError(msg) + schemas = Schemas(field_schemas(context)) + return schemas.document(schemas.of(kind_field(asked))) + + +def field_schemas(context: Context) -> SchemaLeaf: + """The schema leaf that writes each field alias as a field of its kind, as `context` reads one: `field_json_schema`'s.""" + + def leaf(annotation: object, schemas: Schemas) -> JSONSchema | None: + kind = field_kind(annotation) + if kind is None or not is_alias(annotation): + return None + alias = annotation + static = annotation in _STATIC_SIZE + return schemas.defined( + alias, alias.__name__, lambda: _field_schema(kind, static, context, schemas) + ) + + return leaf + + +def _field_schema( + kind: type[Definition[Any]], static: bool, context: Context, schemas: Schemas +) -> JSONSchema: + """A field of `kind` as `context` reads it: one a definition in scope reads, or one none of them claims.""" + table = context.tables.get(kind, {}) + branches: list[JSONValue] = [] + for definition in table.values(): + if static and cast("CodecDefinition[Any]", definition).size != "static": + continue + branches.extend(_read_by(definition, schemas)) + branches.extend(_unclaimed(table)) + return {"anyOf": branches} + + +def written_name(definition: Definition[Any]) -> JSONSchema: + """The JSON Schema of each name a document writes for `definition`: its name, or `r` and a size for raw bits.""" + if isinstance(definition, DataTypeDefinition) and definition.name == RAW_BYTES_NAME: + return {"type": "string", "pattern": _RAW_BYTES_SCHEMA_PATTERN} + return {"const": definition.name} + + +def _read_by(definition: Definition[Any], schemas: Schemas) -> list[JSONValue]: + """The fields `definition` reads: an object of its name and configuration, and its bare name when it needs no configuration. + + Raw bits' name carries their configuration, so what is written beside + it holds nothing. + """ + carried = isinstance(definition, DataTypeDefinition) and definition.name == RAW_BYTES_NAME + name = written_name(definition) + bare = carried or not definition.requires_configuration + envelope: JSONSchema = { + "type": "object", + "properties": { + "name": name, + "configuration": schemas.of( + EmptyConfiguration if carried else definition.configuration + ), + "must_understand": {"const": True}, + }, + "required": ["name"] if bare else ["name", "configuration"], + "additionalProperties": False, + } + return [name, envelope] if bare else [envelope] + + +def _unclaimed(table: Mapping[str, Definition[Any]]) -> list[JSONValue]: + """The fields no definition in `table` claims: a name none of them is written with, bare or with any configuration.""" + claimed: list[JSONValue] = [written_name(definition) for definition in table.values()] + name: JSONSchema = {"type": "string", "pattern": EXTENSION_NAME_SCHEMA_PATTERN} + if len(claimed) != 0: + name["not"] = {"anyOf": claimed} + envelope: JSONSchema = { + "type": "object", + "properties": { + "name": name, + "configuration": {"type": "object"}, + "must_understand": {"const": True}, + }, + "required": ["name"], + "additionalProperties": False, + } + return [name, envelope] + + @dataclass(frozen=True, slots=True) class _Checker: """A TypedDict's checker, compiled once, and whether a value of it can hold a nested field.""" @@ -629,25 +977,57 @@ def leaf(annotation: object) -> Parser | None: def _checked( shape: type, value: object, loc: Loc ) -> tuple[object, Problems, tuple[_NestedField, ...]]: - """`value` checked as `shape`: the typed value, every problem, and the fields nested in it, in order.""" + """`value` checked as `shape`: the typed value, every problem, and the fields nested in it, in order, each put back as it was written.""" + typed, found, nested = _typed(shape, value, loc) + return _put_back(typed, _as_written), found, nested + + +def _typed( + shape: type, value: object, loc: Loc +) -> tuple[object, Problems, tuple[_NestedField, ...]]: + """`value` checked as `shape`, each nested field still where the checker met it: the typed value, every problem, and those fields, in order.""" checker = _checker(shape) typed, found = checker.parse(value, loc) if not checker.nests: return typed, found, () nested: list[_NestedField] = [] - return _put_back(typed, nested), found, tuple(nested) + _collect(typed, nested) + return typed, found, tuple(nested) -def _put_back(value: object, nested: list[_NestedField]) -> object: - """`value` with each nested field the checker handed back put back as its JSON, collected in order.""" +def _collect(value: object, nested: list[_NestedField]) -> None: + """Each nested field the checker handed back in `value`, in order.""" if isinstance(value, _NestedField): nested.append(value) - return value.json - if isinstance(value, tuple): - return tuple(_put_back(entry, nested) for entry in cast("tuple[object, ...]", value)) - if isinstance(value, dict): - entries = cast("dict[str, object]", value) - return {key: _put_back(entry, nested) for key, entry in entries.items()} + elif is_tuple(value): + for entry in value: + _collect(entry, nested) + elif is_object(value): + for entry in value.values(): + _collect(entry, nested) + + +def _as_written(field: _NestedField) -> JSONValue: + """A nested field put back as it was written.""" + return field.json + + +def _declared(field: _NestedField) -> JSONValue: + """A nested field put back as it was written, but for the members its envelope does not declare, which are reported and left out, as the checker leaves out a key a closed TypedDict does not declare.""" + if not isinstance(field.json, Mapping): + return field.json + envelope = cast("Mapping[str, JSONValue]", field.json) + return {key: value for key, value in envelope.items() if key in ENVELOPE_KEYS} + + +def _put_back(value: object, put: Callable[[_NestedField], JSONValue]) -> object: + """`value` with each nested field the checker handed back put back as `put` gives it.""" + if isinstance(value, _NestedField): + return put(value) + if is_tuple(value): + return tuple(_put_back(entry, put) for entry in value) + if is_object(value): + return {key: _put_back(entry, put) for key, entry in value.items()} return value @@ -656,6 +1036,16 @@ def _usable(problems: Sequence[ValidationProblem]) -> bool: return all(found.kind == "unknown_key" for found in problems) +def read_only(configuration: C) -> C: + """`configuration` as a definition's functions are handed it: a read-only view of the field's own, so a function that assigns a member fails there. + + The view is of the members: what a member holds, an object among a + struct's `fields` say, is the field's own, and a function that writes + into one writes into the field. + """ + return cast("C", MappingProxyType(cast("Mapping[str, JSONValue]", configuration))) + + def asked(definition: Definition[Any], what: str, ask: Callable[[], T], at: Loc | None = None) -> T: """What `ask`, a call of `definition`'s `what`, gives. @@ -689,14 +1079,12 @@ def ruled( def _located(prefix: Loc, problems: Iterable[ValidationProblem]) -> Problems: - return tuple( - ValidationProblem((*prefix, *found.loc), found.message, found.kind) for found in problems - ) + return tuple(dataclasses.replace(found, loc=(*prefix, *found.loc)) for found in problems) def _envelope(field: _NestedField) -> Problems: """What is wrong with a nested field's envelope, at the field.""" - return _located(field.loc, envelope_problems(field.json, allow_must_understand_false=False)) + return _located(field.loc, field.kind.envelope_problems(field.json)) def _configuration_checked( @@ -704,17 +1092,22 @@ def _configuration_checked( ) -> tuple[T | None, Problems]: """`value` checked as `shape`, as `typed_json.check` checks it, and each nested field's envelope judged. - The step a definition's `judge` starts from. A member typed with a + The step a definition's `read_configuration` starts from. A member typed with a field alias holds a metadata field, whose envelope is judged as a - document's is: a stray member, or a `must_understand` of `false`, is a - problem of the configuration, and the value does not come back. + document's is: a `must_understand` of `false` is a problem of the + configuration, and the value does not come back; a stray member is an + unknown key, reported and left out, and the value still comes back, + as the checker reports and leaves out a key a closed TypedDict does + not declare. """ refined, problems = refine_json(value, loc) if len(problems) != 0: return None, problems - typed, found, nested = _checked(shape, refined, loc) + typed, found, nested = _typed(shape, refined, loc) problems = (*found, *(problem for field in nested for problem in _envelope(field))) - return (cast("T", typed) if _usable(problems) else None), problems + if not _usable(problems): + return None, problems + return cast("T", _put_back(typed, _declared)), problems def named_configuration( @@ -730,9 +1123,9 @@ def named_configuration( """ if isinstance(value, str): return value, None, () - if not isinstance(value, Mapping): + if not is_object(value): return None, None, () - entry = cast("Mapping[str, object]", value) + entry = value name = entry.get("name") if not isinstance(name, str): return None, None, () @@ -740,51 +1133,338 @@ def named_configuration( return name, None, () configuration = entry["configuration"] if not isinstance(configuration, Mapping): - return name, None, problem(("configuration",), f"expected an object, got {configuration!r}") + return ( + name, + None, + problem(("configuration",), f"expected an object, got {shown(configuration)}"), + ) return name, cast("Mapping[str, object]", configuration), () -Unread = Literal["out_of_scope", "invalid"] -"""A field no definition read: nothing in scope claims its name, or it could not be read.""" - -Resolution = Literal["read"] | Unread -"""What a scope made of a field: read by the definition that claims it, or unread, and why.""" - - def _nothing_nested() -> Nested: - """What a field that holds no field, or was not read, holds inside: nothing.""" + """What a field that holds no field it read holds inside: nothing.""" return {} -@dataclass(frozen=True, slots=True) -class Resolved(Generic[D]): - """One metadata field, as read in a scope: its JSON, and what the scope made of it. +@dataclass(frozen=True, slots=True, kw_only=True) +class AcceptedField(Generic[D]): + """A field a definition in scope read: the name it is written with, the definition, and the configuration it allowed. - `resolution` is what became of the configuration: read by the - definition that claims the name, claimed by nothing, or not readable. A problem with the envelope around it -- a stray member, a - `must_understand` of `false` -- is reported with the field, and leaves - the resolution as it is; so is a problem of a field the configuration - holds, which is that field's own, with its own resolution in `nested`. + `must_understand` of `false` -- is reported with the field and leaves + it read; so is a problem of a field its configuration holds, which is + that field's own, as `nested` says. Two fields are equal when they + read the same, however each was spelled, as `field_key` compares + them: `"bytes"` and `{"name": "bytes"}` are one field, and so are a + blosc with and without the `typesize` that `noshuffle` ignores, which + the definition's `canonical` folds. Equal fields hash alike. """ json: JSONValue - """The field as written, refined: arrays as tuples. `None` for a value that was not JSON, as for `null`.""" - resolution: Resolution - definition: D | None - """The definition that claims the field's name; None when nothing in scope does, or it names none.""" - configuration: Mapping[str, JSONValue] | None - """The configuration, type-checked and allowed by the rules, when the field was read; None otherwise.""" + """The field as written, refined: arrays as tuples; it takes no part in equality.""" + name: str + """The name it is written with: `"r16"`, though its definition is filed under `r*`.""" + definition: D + """The definition that read it.""" + configuration: Mapping[str, JSONValue] + """The configuration, type-checked and allowed by the rules; for raw bits, what the name carries. + + Each field it holds is written as a document writes it, as that + field's `to_json` writes it, so the configuration says what was read + however it was spelled: a shard's `"crc32c"` and `{"name": "crc32c"}` + are one index codec. + """ nested: Nested = dataclasses.field(default_factory=_nothing_nested) """The fields the configuration holds, each as the scope read it, by where it sits in the configuration. A struct's field types at `("fields", 0, "data_type")`, a shard's codecs at `("codecs", 0)`: what a definition's functions consult about - the fields inside its own. Empty unless the field was read. + the fields inside its own. """ + read_as: type[Definition[Any]] = dataclasses.field(init=False, repr=False) + """The kind of metadata it was read as: its definition's.""" + + def __eq__(self, other: object) -> bool: + if not is_field(other): + return NotImplemented + return field_key(self) == field_key(other) + + def __hash__(self) -> int: + return hash(field_key(self)) + + def __post_init__(self) -> None: + # The runtime half of the annotations: a field read by hand, as an + # extension's may be, fails here rather than where a function trusts it. + definition = cast("object", self.definition) + kind = kind_of(cast("Definition[Any]", definition)) + refusal = ( + f"a field read is read by a definition of a kind, got {definition!r}" + if kind is None + else _misread(definition, kind, self.name) + ) + if refusal is not None: + raise TypeError(refusal) + object.__setattr__(self, "read_as", kind) + # What the field hands out is read-only at every level, so a field + # cannot be put in a state its key, `==` and `refines` disagree about. + object.__setattr__(self, "json", frozen(self.json)) + object.__setattr__(self, "configuration", frozen(self.configuration)) + object.__setattr__(self, "nested", MappingProxyType(dict(self.nested))) + + def __reduce__(self) -> tuple[Callable[..., AcceptedField[Any]], tuple[object, ...]]: + # Read-only views do not pickle: the field pickles as what it was built from. + return _accepted_field, ( + copied(self.json), + self.name, + self.definition, + copied(self.configuration), + dict(self.nested), + ) + + def to_json(self) -> JSONValue: + """The field as a document writes it, for every reader: its configuration as read, sharing nothing with the field. + + The envelope takes the fewest words every reader takes: a data type + with nothing to configure is its bare name, as core data types have + been written since Zarr v3.0; any other field is an object, + `{"name": ...}`, since a Zarr v3.0 reader takes no bare name in + `codecs` + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L585-L592). + A name that carries its configuration, as raw bits' does, is + written alone. + """ + return copied(written_json(self)) + + +@dataclass(frozen=True, slots=True, kw_only=True) +class UnclaimedField: + """A field nothing in scope claims: an extension the scope leaves unjudged, which is what keeps the format open. + + Equal to another when it is written with the same name and + configuration, however each was spelled: nothing in scope interprets + its configuration, so it compares as JSON text, as `field_key` says. + """ + + json: JSONValue + """The field as written, refined: arrays as tuples; it takes no part in equality.""" + name: str + """The name nothing in scope claims.""" + read_as: type[Definition[Any]] + """The kind of metadata it was read as: what a definition that claimed it would be.""" + configuration: Mapping[str, JSONValue] = dataclasses.field(init=False) + """The configuration as written, which nothing judged; empty when none is written.""" + + def __eq__(self, other: object) -> bool: + if not is_field(other): + return NotImplemented + return field_key(self) == field_key(other) + + def __hash__(self) -> int: + return hash(field_key(self)) + + def __post_init__(self) -> None: + # The runtime half of the annotations; `read_as` with its type + # arguments dropped, as `resolve` drops them. + object.__setattr__(self, "read_as", as_kind(self.read_as)) + name = cast("object", self.name) + if not isinstance(name, str) or not self.read_as.well_named(name): + msg = ( + "a field nothing in scope claims is named as a document names a " + f"{self.read_as.label}, got {name!r}" + ) + raise TypeError(msg) + _, written, _ = self.read_as.named_configuration(self.json) + configuration: Mapping[str, object] = {} if written is None else written + object.__setattr__(self, "json", frozen(self.json)) + object.__setattr__( + self, "configuration", frozen(cast("Mapping[str, JSONValue]", configuration)) + ) + + def __reduce__(self) -> tuple[Callable[..., UnclaimedField], tuple[object, ...]]: + # Read-only views do not pickle: the field pickles as what it was built from. + return _unclaimed_field, (copied(self.json), self.name, self.read_as) + @property + def definition(self) -> None: + """The definition that read it: none did.""" + return None + + @property + def nested(self) -> Nested: + """The fields its configuration holds as the scope read them: none, since nothing read its configuration.""" + return _nothing_nested() + + def to_json(self) -> JSONValue: + """The field as a document writes it, sharing nothing with the field: its configuration as written, in the envelope every reader takes, as `AcceptedField.to_json` writes one.""" + return copied(written_json(self)) + + +@dataclass(frozen=True, slots=True, kw_only=True) +class RefusedField(Generic[D]): + """A field that could not be read -- not a field at all, not JSON, or refused by the definition that claims its name -- as its problems say.""" + + json: JSONValue | UNSET + """The field as written, refined: arrays as tuples; `UNSET` when it is not JSON, which no document holds.""" + name: str | None + """The name it is written with; None when it names none.""" + read_as: type[Definition[Any]] + """The kind of metadata it was read as.""" + definition: D | None = None + """The definition that claims its name and refused it; None when nothing in scope claims it, or it names none.""" + nested: Nested = dataclasses.field(default_factory=_nothing_nested) + """The fields its configuration holds, each as the scope read it; empty when its configuration was not checked against its TypedDict.""" + + def __eq__(self, other: object) -> bool: + if not is_field(other): + return NotImplemented + return field_key(self) == field_key(other) + + def __hash__(self) -> int: + return hash(field_key(self)) + + def __post_init__(self) -> None: + # The runtime half of the annotations; `read_as` with its type + # arguments dropped, as `resolve` drops them. + kind = as_kind(self.read_as) + object.__setattr__(self, "read_as", kind) + definition = cast("object", self.definition) + refusal = None if definition is None else _misread(definition, kind, self.name) + if refusal is not None: + raise TypeError(refusal) + if self.json is not UNSET: + object.__setattr__(self, "json", frozen(self.json)) + object.__setattr__(self, "nested", MappingProxyType(dict(self.nested))) + + def __reduce__(self) -> tuple[Callable[..., RefusedField[Any]], tuple[object, ...]]: + # Read-only views do not pickle: the field pickles as what it was built from. + return _refused_field, ( + UNSET if self.json is UNSET else copied(self.json), + self.name, + self.read_as, + self.definition, + dict(self.nested), + ) + + +ResolvedField = TypeAliasType( + "ResolvedField", "AcceptedField[D] | UnclaimedField | RefusedField[D]", type_params=(D,) +) +"""One metadata field as a scope read it: read by the definition that claims its name, claimed by nothing, or refused.""" + + +def _accepted_field( + json: JSONValue, + name: str, + definition: Definition[Any], + configuration: Mapping[str, JSONValue], + nested: Nested, +) -> AcceptedField[Any]: + """An accepted field built again from what it pickled as.""" + return AcceptedField( + json=json, name=name, definition=definition, configuration=configuration, nested=nested + ) + + +def _unclaimed_field(json: JSONValue, name: str, read_as: type[Definition[Any]]) -> UnclaimedField: + """An unclaimed field built again from what it pickled as.""" + return UnclaimedField(json=json, name=name, read_as=read_as) + + +def _refused_field( + json: JSONValue | UNSET, + name: str | None, + read_as: type[Definition[Any]], + definition: Definition[Any] | None, + nested: Nested, +) -> RefusedField[Any]: + """A refused field built again from what it pickled as.""" + return RefusedField(json=json, name=name, read_as=read_as, definition=definition, nested=nested) + + +_FIELDS: Final = (AcceptedField, UnclaimedField, RefusedField) +"""The three things a scope makes of a field.""" + + +def _misread(definition: object, kind: type[Definition[Any]], name: object) -> str | None: + """What is wrong with `definition` as the one that read a field named `name` as a `kind`; None when nothing is.""" + if not isinstance(definition, kind): + return f"a field read as a {kind.__name__} is read by one, got {definition!r}" + filed = cast("Definition[Any]", definition).name + if not isinstance(name, str) or spelled(kind, name)[0] != filed: + return f"a field named {name!r} is read by the definition filed under it, got {filed!r}" + return None + + +def field_key(field: ResolvedField[Any]) -> tuple[object, ...]: + """What `==` and `hash` compare of a field: what it means, not how it was spelled. + + A field read compares by the definition that read it, its + configuration in the canonical spelling the definition gives it -- what + `canonical_of` spells -- as JSON text, and the fields it holds, each by + its own key. A field nothing claims compares by its name and its + configuration as written, as JSON text: nothing interprets it. A field + refused compares by what was written, and by what refused it. So two + fields are one when what a reader understands of them reads the same, + and what none interprets is written alike. + """ + if isinstance(field, AcceptedField): + definition = cast("Definition[Any]", field.definition) + configuration: JSONValue = dict(field.configuration) + # A field it holds is compared by its own key, so its place holds + # nothing before the definition's `canonical` sees the rest, as + # `_canonical_field` orders it: `canonical` folds only the + # definition's own members. + for loc in field.nested: + configuration = _replaced(configuration, loc, None) + view = read_only(cast("Mapping[str, JSONValue]", configuration)) + spelled = asked( + definition, + "canonical", + lambda: json_text(cast("JSONValue", dict(definition.canonical(view)))), + ) + return ("read", definition, spelled, _nested_key(field.nested)) + if isinstance(field, UnclaimedField): + return ("unclaimed", field.read_as, field.name, json_text(field.configuration)) + return ( + "refused", + field.read_as, + field.name, + field.definition, + UNSET if field.json is UNSET else json_text(field.json), + _nested_key(field.nested), + ) -Nested: TypeAlias = Mapping[Loc, Resolved[Any]] + +def own_key(field: AcceptedField[Any]) -> tuple[object, ...]: + """What `field_key` compares of a read field without the fields it holds: the definition and the canonical spelling of its own members.""" + return field_key(field)[:3] + + +def _nested_key(nested: Nested) -> tuple[tuple[Loc, tuple[object, ...]], ...]: + """The fields a configuration holds, each by its key, where it sits.""" + return tuple((loc, field_key(inner)) for loc, inner in nested.items()) + + +def document_json(field: ResolvedField[Any]) -> JSONValue | UNSET: + """A field as a document writes it, holding the field's own values: as `to_json` writes it, or as it was written when it was refused, `UNSET` for one that was not JSON. + + What a writer serializes, which changes nothing, so it copies nothing; + `to_json` is this, copied. + """ + if isinstance(field, RefusedField): + return field.json + return written_json(field) + + +def written_json(field: AcceptedField[Any] | UnclaimedField) -> JSONValue: + """A field a scope read or left unclaimed, as a document writes it: `document_json` of one that is JSON.""" + kind = field.read_as + if isinstance(field, AcceptedField) and kind.spelled(field.name)[1] is not None: + return kind.envelope_json(field.name, {}) + return kind.envelope_json(field.name, field.configuration) + + +Nested: TypeAlias = Mapping[Loc, ResolvedField[Any]] """The fields a configuration holds, each as the scope read it, by where it sits in the configuration.""" @@ -806,8 +1486,8 @@ class Chunk: lengths: Lengths | None = None """Per axis, the lengths the chunks take along it; None when not even the number of axes is known.""" - data_type: Resolved[DataTypeDefinition[Any]] | None = None - """The data type field of the values, as a scope read it; None when no field says what they are.""" + data_type: ResolvedField[DataTypeDefinition[Any]] | None = None + """The data type field of the values, as a scope read it; None when no field says what they are: a document naming none, which its reading holds as `UNSET`, hands the pipeline a chunk of no known type.""" def __post_init__(self) -> None: lengths = cast("object", self.lengths) @@ -825,19 +1505,29 @@ def rank(self) -> int | None: return None if self.lengths is None else len(self.lengths) +def is_field(value: object) -> TypeGuard[ResolvedField[Any]]: + """Whether `value` is a field a scope read: `AcceptedField`, `UnclaimedField` or `RefusedField`.""" + return isinstance(value, _FIELDS) + + +def held(field: ResolvedField[D] | UNSET) -> AcceptedField[D] | UnclaimedField: + """`field`, one a read found nothing wrong with: `AcceptedField` or `UnclaimedField`; `TypeError` for one refused or never read, which such a read rules out.""" + if isinstance(field, (AcceptedField, UnclaimedField)): + return field + msg = f"expected a field a read found nothing wrong with, got {field!r}" + raise TypeError(msg) + + def _is_data_type_field(value: object) -> bool: - """Whether `value` is a data type field a scope read: one read as a data type, or by nothing.""" - if not isinstance(value, Resolved): - return False - definition = cast("Resolved[Any]", value).definition - return definition is None or isinstance(definition, DataTypeDefinition) + """Whether `value` is a field a scope read as a data type.""" + return is_field(value) and value.read_as is DataTypeDefinition def _is_lengths(value: object) -> TypeGuard[Lengths]: """Whether `value` is chunk lengths: per axis, a frozenset of integers, or None.""" - if not isinstance(value, tuple): + if not is_tuple(value): return False - for axis in cast("tuple[object, ...]", value): + for axis in value: if axis is None: continue if not isinstance(axis, frozenset) or not all( @@ -848,27 +1538,70 @@ def _is_lengths(value: object) -> TypeGuard[Lengths]: return True -def configuration_of(resolved: Resolved[Any], definition: Definition[C]) -> C | None: +def fields_of( + resolved: ResolvedField[Any], loc: Loc = () +) -> Iterator[tuple[Loc, ResolvedField[Any]]]: + """`resolved`, a field a scope read, where it sits, then each field it holds and theirs in turn, each where it sits. + + `loc` is where `resolved` sits; a field it holds sits in its + configuration, at the kind's `configuration_loc` and its place, as `resolve` + locates its problems. What each holds is its `nested`: a field whose + configuration was not checked holds none. + """ + yield loc, resolved + for place, inner in resolved.nested.items(): + yield from fields_of(inner, (*resolved.read_as.configuration_loc(loc), *place)) + + +def with_problems( + fields: Iterable[tuple[Loc, ResolvedField[Any]]], problems: Sequence[ValidationProblem] +) -> Iterator[tuple[Loc, ResolvedField[Any], Problems]]: + """Each of `fields`, with where it sits, and the problems among `problems` located in it, in the fields it holds too. + + `fields` and `problems` are one read's: a reading's `fields()` and + `problems`, or `fields_of` a field and the problems `resolve` gave + with it. A field's problems are those it was read with, and those the + document found with it where it stands -- its place in the pipeline, + the chunk it is handed, the array's shape -- so a field with none is + valid there, and `canonical_of` spells it. A function of the problems, + as zod's `treeifyError` is of the issues, grouping them at every + depth: a problem with a shard's inner codec is the inner codec's, and + the shard's. Each field comes before the fields it holds, as + `fields_of` gives them, so the last field whose problems hold a + problem is the innermost field holding it. A problem in no field -- + with the fill value, with the shape -- is in none's. + """ + located = list(fields) + held: dict[Loc, list[ValidationProblem]] = {loc: [] for loc, _ in located} + for found in problems: + for depth in range(len(found.loc) + 1): + holder = held.get(found.loc[:depth]) + if holder is not None: + holder.append(found) + for loc, field in located: + yield loc, field, tuple(held[loc]) + + +def configuration_of(resolved: ResolvedField[Any], definition: Definition[C]) -> C | None: """The configuration `resolved` holds, typed as `definition` declares it, if `definition` read it. - `Resolved` holds a configuration as the mapping every one is; asked - with the definition that read the field, this is the same mapping, as - its TypedDict. None when another definition read it, or none did. + `AcceptedField` holds a configuration as the mapping every one is; asked with + the definition that read the field -- or one equal to it, as a + pickled field's is -- this is the same mapping, as its TypedDict. None + when another definition read it, or none did. """ - if resolved.definition is not definition or resolved.configuration is None: + if not isinstance(resolved, AcceptedField) or resolved.definition != definition: return None return cast("C", resolved.configuration) -def fill_value_problems( - data_type: Resolved[DataTypeDefinition[Any]], value: object, loc: Loc = () -) -> Problems: +def fill_value_problems(data_type: ResolvedField[F], value: object, loc: Loc = ()) -> Problems: """What is wrong with `value` as a fill value of `data_type`, a data type field a scope read. `value` is refined to JSON first: not JSON is the first verdict, whatever the data type. It is then checked against the JSON shape the data type's definition declares, and judged by its fill value rules, as - `judge` judges a configuration: a key the shape does not declare is + `read_configuration` reads a configuration: a key the shape does not declare is reported and left out, and the rules still judge the rest. The rules see the fields the configuration holds as the scope read them: a struct judges each field's fill value by that field's own type. A data @@ -876,22 +1609,67 @@ def fill_value_problems( value unjudged. `loc` prefixes every problem. """ refined, problems = refine_json(value, loc) - definition = data_type.definition - configuration = data_type.configuration - if len(problems) != 0 or definition is None or configuration is None: - return problems + if len(problems) != 0 or not isinstance(data_type, AcceptedField): + return with_input(problems, value, loc) + definition, configuration = data_type.definition, data_type.configuration typed, problems = _fill_value_parser(definition.fill_value)(refined, loc) if not _usable(problems): - return problems + return with_input(problems, value, loc) refused = ruled( definition, - lambda: definition.fill_value_rules(configuration, data_type.nested, typed), + lambda: definition.fill_value_rules(read_only(configuration), data_type.nested, typed), loc, ) - return (*problems, *refused) + return with_input((*problems, *refused), value, loc) + + +def canonical_fill_value(data_type: ResolvedField[F], value: object) -> JSONValue | UNSET: + """`value`, a fill value of `data_type`, a data type field a scope read, in the one spelling its value has; `UNSET` when it has a problem. + + As the data type's `fill_value_canonical` spells it, so two fill + values of a data type are one value exactly when their canonical + spellings are written alike -- the same JSON, as `json.dumps` writes + it, which `==` is not: it takes `-0.0` for `0.0`. A fill value + `fill_value_problems` finds a problem with has no canonical spelling, + as `canonical_of` gives a field with a problem none: `UNSET`, since + `None` is the JSON `null`, a fill value of a data type the scope did + not read, which spells a fill value as written. + """ + if len(fill_value_problems(data_type, value)) != 0: + return UNSET + refined, _ = refine_json(value, ()) + return spelled_canonically(data_type, refined) + + +def spelled_canonically(data_type: ResolvedField[F], value: JSONValue) -> JSONValue: + """`value`, a fill value of `data_type` with no problem, in its canonical spelling, as `canonical_fill_value` gives it, without judging it again. + + Its `fill_value_canonical` is the extension author's code: what it + gives is checked to be JSON, and an error it raises says which data + type's canonical spelling raised it. + """ + if not isinstance(data_type, AcceptedField): + return value + definition, configuration = data_type.definition, data_type.configuration + spelled = asked( + definition, + "fill_value_canonical", + lambda: cast( + "object", + definition.fill_value_canonical(read_only(configuration), data_type.nested, value), + ), + ) + refined, problems = refine_json(spelled, ()) + if len(problems) != 0: + msg = ( + f"{definition.name!r}: its fill_value_canonical gives JSON, got {spelled!r}: " + f"{problems[0].message}" + ) + raise TypeError(msg) + return refined -def storage_of(data_type: Resolved[DataTypeDefinition[Any]]) -> StorageClass | None: +def storage_of(data_type: ResolvedField[DataTypeDefinition[Any]]) -> StorageClass | None: """How the values of `data_type`, a data type field a scope read, are stored; None when unknown. Unknown when the scope did not read it, or its definition does not @@ -899,14 +1677,13 @@ def storage_of(data_type: Resolved[DataTypeDefinition[Any]]) -> StorageClass | N checked to be a storage class, and an error it raises says which data type's storage raised it. """ - definition = data_type.definition - configuration = data_type.configuration - if definition is None or configuration is None: + if not isinstance(data_type, AcceptedField): return None + definition, configuration = data_type.definition, data_type.configuration found = asked( definition, "storage", - lambda: cast("object", definition.storage(configuration, data_type.nested)), + lambda: cast("object", definition.storage(read_only(configuration), data_type.nested)), ) if found is not None and found not in get_args(StorageClass): msg = ( @@ -918,7 +1695,7 @@ def storage_of(data_type: Resolved[DataTypeDefinition[Any]]) -> StorageClass | N def chunk_grid_lengths( - chunk_grid: Resolved[ChunkGridDefinition[Any]], shape: tuple[int, ...], loc: Loc = () + chunk_grid: ResolvedField[ChunkGridDefinition[Any]], shape: tuple[int, ...], loc: Loc = () ) -> tuple[Lengths, Problems]: """The lengths the chunks of `chunk_grid`, a chunk grid field a scope read, take along each axis of an array of `shape`, and what is wrong with the grid over it. @@ -936,20 +1713,23 @@ def chunk_grid_lengths( definition, not the field. """ unknown: Lengths = (None,) * len(shape) - definition = chunk_grid.definition - configuration = chunk_grid.configuration - if definition is None or configuration is None: + if not isinstance(chunk_grid, AcceptedField): return unknown, () + definition, configuration = chunk_grid.definition, chunk_grid.configuration at = (*loc, "configuration") problems = ruled( - definition, lambda: definition.shape_rules(configuration, chunk_grid.nested, shape), at + definition, + lambda: definition.shape_rules(read_only(configuration), chunk_grid.nested, shape), + at, ) if len(problems) != 0: - return unknown, problems + return unknown, with_input(problems, chunk_grid.json, loc) lengths = asked( definition, "chunk lengths", - lambda: cast("object", definition.chunk_lengths(configuration, chunk_grid.nested, shape)), + lambda: cast( + "object", definition.chunk_lengths(read_only(configuration), chunk_grid.nested, shape) + ), at, ) if not _is_lengths(lengths): @@ -968,8 +1748,8 @@ def chunk_grid_lengths( def resolve( - data: object, kind: type[D], context: Context, loc: Loc = () -) -> tuple[Resolved[D], Problems]: + data: object, kind: type[D], context: Scope, loc: Loc = () +) -> tuple[ResolvedField[D], Problems]: """`data`, one metadata field, read as a `kind` in `context`: what the scope made of it, and every problem. All three steps for one field. `data` is refined to JSON and its @@ -979,91 +1759,134 @@ def resolve( configuration is checked against its TypedDict and judged by its rules; each nested field the check met is read the same way, in the same scope, and what is wrong with one is its own, reported where it - sits, as with a document's fields. A name nothing claims is - `out_of_scope`: an unmodelled - extension, left unjudged, which is what keeps the format open. `loc` - prefixes every problem. `kind` is one of `KINDS`, with or without - type arguments; anything else is a `TypeError`. + sits, as with a document's fields. What comes back is `AcceptedField` by the + definition that claims the name; `UnclaimedField` when nothing in scope + claims it, an unmodelled extension left unjudged, which is what keeps + the format open; or `RefusedField`, with the problems that say why. `loc` + prefixes every problem. `kind` is one of the five kinds -- + `CodecDefinition`, `DataTypeDefinition`, `ChunkGridDefinition`, + `ChunkKeyEncodingDefinition`, `StorageTransformerDefinition` -- with + or without type arguments; anything else is a `TypeError`. """ asked = as_kind(kind) + if context.format is not None and context.format != format_of(asked): + msg = ( + f"a {asked.label} is a field of a Zarr v{format_of(asked)} document, read in a scope " + f"of that format or of none, got a v{context.format} scope" + ) + raise TypeError(msg) refined, problems = refine_json(data, loc) if len(problems) != 0: - return Resolved(None, "invalid", None, None), problems + # Not JSON, so not read; its name, if it has one, still says what + # claims it, and one the spec does not give an extension is a + # problem here as on the other path, asked of no definition. + name = asked.named_configuration(data)[0] + bad = None if name is None else asked.name_problem(name, asked.name_loc(loc)) + claimant = None if name is None or bad is not None else context.claimant(asked, name) + refused = RefusedField(json=UNSET, name=name, read_as=asked, definition=claimant) + found = problems if bad is None else (bad, *problems) + return cast("ResolvedField[D]", refused), with_input(found, data, loc) resolved, found = _resolve_field(refined, asked, context, loc) - return cast("Resolved[D]", resolved), found + return cast("ResolvedField[D]", resolved), with_input(found, data, loc) def _resolve_field( - data: JSONValue, kind: type[Definition[Any]], context: Context, loc: Loc -) -> tuple[Resolved[Definition[Any]], Problems]: + data: JSONValue, kind: type[Definition[Any]], context: Scope, loc: Loc +) -> tuple[ResolvedField[Definition[Any]], Problems]: """A refined field with its envelope judged, then read. - The resolution is what became of the configuration. A stray member or - a `must_understand` of `false` says nothing about it, so it is reported - beside the field that was read, which later layers can still judge. + What the scope made of it is what became of the configuration. A + stray member or a `must_understand` of `false` says nothing about it, + so it is reported beside the field that was read, which later layers + can still judge. """ - envelope = _located(loc, envelope_problems(data, allow_must_understand_false=False)) + envelope = _located(loc, kind.envelope_problems(data)) resolved, found = _read(data, kind, context, loc) return resolved, (*envelope, *found) def _read( - data: JSONValue, kind: type[Definition[Any]], context: Context, loc: Loc -) -> tuple[Resolved[Definition[Any]], Problems]: - name, given, malformed = named_configuration(data) + data: JSONValue, kind: type[Definition[Any]], context: Scope, loc: Loc +) -> tuple[ResolvedField[Definition[Any]], Problems]: + name, given, malformed = kind.named_configuration(data) if name is None: - return Resolved(data, "invalid", None, None), () + return RefusedField(json=data, name=None, read_as=kind), () + if not kind.well_named(name): + # The envelope rule every reader runs first reports it; no + # definition is asked to claim it. + return RefusedField(json=data, name=name, read_as=kind), () definition = context.claimant(kind, name) if len(malformed) != 0: # A configuration that is not an object, which the envelope's # problems say; the name still says what claims the field. - return Resolved(data, "invalid", definition, None), () + return RefusedField(json=data, name=name, read_as=kind, definition=definition), () if definition is None: - return Resolved(data, "out_of_scope", None, None), () - _, carried = spelled(kind, name) + return UnclaimedField(json=data, name=name, read_as=kind), () + _, carried = kind.spelled(name) if carried is not None: - return _read_carried(data, definition, given, carried, loc) - at = (*loc, "configuration") + return _read_carried(data, name, kind, definition, given, carried, loc) + at = kind.configuration_loc(loc) if given is None and definition.requires_configuration: missing = problem(at, f"{name!r} requires a configuration", "missing_key") - return Resolved(data, "invalid", definition, None), missing - typed, found, nested = _checked(definition.configuration, {} if given is None else given, at) - # The rules may read a field the configuration holds by its name, so - # they are asked only when each one is named; any other problem with - # one is its own, reported where it sits, as a document's fields are. - sound = _usable(found) and all(_named(field) for field in nested) - configuration = cast("Mapping[str, JSONValue]", typed) if sound else None - # The fields it holds are read first, so the rules see them as the - # scope read them; their problems are reported after the rules'. - within: dict[Loc, Resolved[Any]] = {} + return RefusedField(json=data, name=name, read_as=kind, definition=definition), missing + typed, found, nested = _typed(definition.configuration, {} if given is None else given, at) + # The fields it holds are read first, each a frame deeper than this + # one, so the rules see them as the scope read them; each is put back + # as a document writes it, so the configuration says what was read + # however each was spelled. Their problems are reported after the + # rules'. + within: dict[Loc, ResolvedField[Any]] = {} + written: dict[Loc, JSONValue] = {} inside: list[ValidationProblem] = [] for field in nested: inside.extend(_envelope(field)) inner, found_inside = _read(field.json, field.kind, context, field.loc) within[field.loc[len(at) :]] = inner + # Refined JSON came in, so what was read of it is JSON, or a field + # refused for what it holds, never for not being JSON. + written[field.loc] = cast("JSONValue", document_json(inner)) inside.extend(found_inside) - inside.extend(_sized(field, inner.definition)) + inside.extend(_sized(field, inner)) + # The rules may read a field the configuration holds by its name, so + # they are asked only when each one is named; any other problem with + # one is its own, reported where it sits, as a document's fields are. + sound = _usable(found) and all(_named(field) for field in nested) + configuration = ( + cast("Mapping[str, JSONValue]", _put_back(typed, lambda field: written[field.loc])) + if sound + else None + ) own = list(found) if configuration is not None: - own.extend(ruled(definition, lambda: definition.rules(configuration, within), at)) + own.extend( + ruled(definition, lambda: definition.rules(read_only(configuration), within), at) + ) if configuration is None or not _usable(own): - return Resolved(data, "invalid", definition, None), (*own, *inside) - return Resolved(data, "read", definition, configuration, within), (*own, *inside) + refused = RefusedField( + json=data, name=name, read_as=kind, definition=definition, nested=within + ) + return refused, (*own, *inside) + read = AcceptedField( + json=data, name=name, definition=definition, configuration=configuration, nested=within + ) + return read, (*own, *inside) def _named(field: _NestedField) -> bool: """Whether a field a configuration holds is named, with an object for its configuration if it has one.""" - name, _, malformed = named_configuration(field.json) + name, _, malformed = field.kind.named_configuration(field.json) return name is not None and len(malformed) == 0 def _read_carried( data: JSONValue, + name: str, + kind: type[Definition[Any]], definition: Definition[Any], given: Mapping[str, object] | None, carried: Mapping[str, JSONValue], loc: Loc, -) -> tuple[Resolved[Definition[Any]], Problems]: +) -> tuple[ResolvedField[Definition[Any]], Problems]: """A field whose name carries its configuration -- raw bits, `r16` -- read by the definition its name is filed under. The document wrote a name, so what is wrong with what the name carries @@ -1072,35 +1895,50 @@ def _read_carried( member of one is a key nothing declares. """ _, beside, _ = _checked( - EmptyConfiguration, {} if given is None else given, (*loc, "configuration") + EmptyConfiguration, + {} if given is None else given, + kind.configuration_loc(loc), ) - configuration, judged = definition.judge(carried) - problems = (*beside, *(ValidationProblem(loc, found.message, found.kind) for found in judged)) - if configuration is None or not _usable(problems): - return Resolved(data, "invalid", definition, None), problems - return Resolved(data, "read", definition, configuration), problems - - -def _sized(field: _NestedField, definition: Definition[Any] | None) -> Problems: + configuration, judged = definition.read_configuration(carried) + # What is wrong with what the name carries is the field's: found at + # the field, where the name is what is there, and what was expected + # of a member of the configuration is not expected of it. + problems = ( + *beside, + *(dataclasses.replace(found, loc=loc, input=UNSET, ctx={}) for found in judged), + ) + # A name is one word: what it carries that the definition does not + # declare cannot be left out, as a stray key beside the name can, so + # anything wrong with what the name carries refuses the field. + if configuration is None or not _usable(beside) or len(judged) != 0: + refused = RefusedField(json=data, name=name, read_as=kind, definition=definition) + return refused, problems + return AcceptedField( + json=data, name=name, definition=definition, configuration=configuration + ), problems + + +def _sized(field: _NestedField, inner: ResolvedField[Any]) -> Problems: """A codec of dynamic size in a member that takes codecs of static size, as a problem at the field. A name nothing in scope claims is left unjudged, its size unknown, as everything else about it is. """ + definition = inner.definition if not field.static or not isinstance(definition, CodecDefinition): return () if definition.size == "static": return () - name, _, _ = named_configuration(field.json) return problem( field.loc, - f"{name!r} is a codec of dynamic size, and only codecs of static size may be used here", + f"{inner.name!r} is a codec of dynamic size, and only codecs of static size may be used " + "here", "invalid_value", ) def canonicalize( - data: object, kind: type[D], context: Context, loc: Loc = () + data: object, kind: type[D], context: Scope, loc: Loc = () ) -> tuple[JSONValue | None, Problems]: """`data`, one metadata field, in its simplest equivalent spelling, and every problem. @@ -1113,47 +1951,75 @@ def canonicalize( `canonical` has the rest -- judged again, so a `canonical` that gives a configuration that does not hold is a `ValueError`, a fault in the definition rather than the field. The envelope takes the fewest - words: the bare name when nothing is configured, and no - `must_understand`, since `true` is what absence means. A name nothing - in scope claims comes back as written, since what it simplifies to is - its own definition's call. + words every reader takes: a data type with nothing to configure is its + bare name, as core data types have been written since Zarr v3.0; any + other field is an object, `{"name": ...}`, since a Zarr v3.0 reader + takes no short-hand name in `codecs` + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L585-L592); + and there is no `must_understand`, since `true` is what absence means. + A name nothing in scope claims keeps the configuration it was written + with, since what it simplifies to is its own definition's call. A + field a scope has read already is spelled so by `canonical_of`, + given its problems. """ resolved, problems = resolve(data, kind, context, loc) - if len(problems) != 0: - return None, problems - if resolved.definition is None: - return resolved.json, () - return _canonical_field(resolved.definition, resolved), () + return canonical_of(resolved, problems), problems -def _canonical_field(definition: Definition[Any], resolved: Resolved[Any]) -> JSONValue: - """A field that read, in its simplest equivalent spelling: the fields it holds first, then its own members. +def canonical_of( + resolved: ResolvedField[Any], problems: Sequence[ValidationProblem] +) -> JSONValue | None: + """`resolved`, a field a scope read, in its simplest equivalent spelling, as `canonicalize` spells one; None when it has a problem. - Each field it holds is spelled from what `resolve` read of it, kept in - `nested`; one nothing in scope claims keeps the spelling it was written - in. + `problems` are the field's, as `resolve` gives them, or + `with_problems` gives each field of a reading. Only a field with none + has a simplest spelling: a simpler spelling of one with a problem + would erase what its author wrote -- a key its TypedDict does not + declare -- or spell what does not hold. What is spelled is what the + scope read, without reading the field again. """ - name, _, _ = named_configuration(resolved.json) - configuration: JSONValue = dict(resolved.configuration or {}) + if len(problems) != 0: + return None + return _simplest(resolved) + + +def _simplest(field: ResolvedField[Any]) -> JSONValue | None: + """A field in its simplest spelling; None when it, or a field it holds, was refused, which has none.""" + if isinstance(field, AcceptedField): + return _canonical_field(field) + if isinstance(field, UnclaimedField): + return field.to_json() + return None + + +def _canonical_field(resolved: AcceptedField[Any]) -> JSONValue | None: + """A field that read, in its simplest equivalent spelling: the fields it holds first, then its own members; None when one it holds was refused.""" + definition, name = resolved.definition, resolved.name + configuration: JSONValue = dict(resolved.configuration) for loc, inner in resolved.nested.items(): - simplest = ( - inner.json if inner.definition is None else _canonical_field(inner.definition, inner) - ) + simplest = _simplest(inner) + if simplest is None: + return None configuration = _replaced(configuration, loc, simplest) - simplified = cast("Mapping[str, JSONValue]", definition.canonical(configuration)) - _, refused = definition.judge(simplified) + # The view every function of a definition is handed; what `canonical` + # gives, the view itself when nothing is folded, is taken as a dict. + view = read_only(cast("Mapping[str, JSONValue]", configuration)) + simplified = asked( + definition, + "canonical", + lambda: dict(cast("Mapping[str, JSONValue]", definition.canonical(view))), + ) + _, refused = definition.read_configuration(simplified) if len(refused) != 0: msg = ( f"{definition.name!r}: its canonical gave {simplified!r}, which does not hold: " f"{list(refused)!r}" ) raise ValueError(msg) - carrying = _carrying_name(definition, simplified) + carrying = definition.carrying_name(simplified) if carrying is not None: return carrying - if len(simplified) == 0: - return name - return {"name": name, "configuration": simplified} + return resolved.read_as.envelope_json(name, simplified) def _replaced(value: JSONValue, path: Loc, new: JSONValue) -> JSONValue: @@ -1161,17 +2027,19 @@ def _replaced(value: JSONValue, path: Loc, new: JSONValue) -> JSONValue: if len(path) == 0: return new step, rest = path[0], path[1:] - if isinstance(step, str): - members = cast("Mapping[str, JSONValue]", value) - return {**members, step: _replaced(members[step], rest, new)} - entries = cast("tuple[JSONValue, ...]", value) - return (*entries[:step], _replaced(entries[step], rest, new), *entries[step + 1 :]) + if isinstance(step, str) and isinstance(value, Mapping): + return {**value, step: _replaced(value[step], rest, new)} + if isinstance(step, int) and isinstance(value, tuple): + return (*value[:step], _replaced(value[step], rest, new), *value[step + 1 :]) + msg = f"{path!r} does not lead into {value!r}" + raise TypeError(msg) __all__ = [ "KINDS", "RAW_BYTES_NAME", "RAW_BYTES_NAME_PATTERN", + "AcceptedField", "Chunk", "ChunkGridDefinition", "ChunkGridField", @@ -1187,32 +2055,45 @@ def _replaced(value: JSONValue, path: Loc, new: JSONValue) -> JSONValue: "EmptyConfiguration", "Lengths", "Nested", - "Resolution", - "Resolved", + "RefusedField", + "ResolvedField", "StaticCodecField", "StorageClass", "StorageTransformerDefinition", "StorageTransformerField", - "Unread", + "UnclaimedField", "as_kind", "asked", + "canonical_fill_value", + "canonical_of", "canonicalize", "chunk_grid_lengths", "configuration_of", + "field_json_schema", + "field_key", + "field_kind", + "field_schemas", + "fill_value_as_written", "fill_value_problems", "kind_of", "multi_byte", "named_configuration", "no_pipelines", "no_rules", + "own_key", + "read_only", "resolve", "ruled", "single_byte", "spelled", + "spelled_canonically", "storage_of", "unchanged", "unknown_chunk", "unknown_lengths", "unknown_storage", "variable_length", + "well_named", + "with_problems", + "written_name", ] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/_hierarchy.py b/packages/zarr-metadata/src/zarr_metadata/v3/_hierarchy.py new file mode 100644 index 0000000000..ac8efbd897 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v3/_hierarchy.py @@ -0,0 +1,287 @@ +"""A Zarr v3 hierarchy: the names and paths of its nodes, and the tree they make. + +`NodeName` and `NodePath` are modeled on zarrs' types of those names, and +hold the spec's rules, which reserve `zarr.json` too. `hierarchy_problems` +judges the node type of each node of a hierarchy, by its path, as a tree. + +Every problem here is one per value, and says its reasons at once, so +what is reported stays proportional to what was read: a path of a +thousand bad names is one problem, not a thousand each repeating the +path. + +Private: consumers import `NodeName` and `NodePath` from `zarr_metadata.v3`, +and the validators from `zarr_metadata.model`. +""" + +from __future__ import annotations + +from typing import TYPE_CHECKING, Literal, NewType, TypeGuard, cast + +from zarr_metadata._json import MetadataValidationError, ValidationProblem, shown, with_input + +if TYPE_CHECKING: + from collections.abc import Mapping, Sequence + +NodeName = NewType("NodeName", str) +"""The name of a node in a Zarr v3 hierarchy. + +The root's is `""`. Any other is not empty, holds no `/`, is not periods +alone -- `.`, `..` -- does not start with the reserved `__`, and is not +`zarr.json`. Case matters: `foo` and `FOO` are two names +(https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L818-L837). +""" + +NodePath = NewType("NodePath", str) +"""The path of a node in a Zarr v3 hierarchy. + +The root's is `/`. Any other's is its parent's path, a `/` unless the +parent is the root, and its name: so a path starts with `/`, does not end +with one, and holds a node name between each two +(https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L211-L229). +""" + +NodeType = Literal["array", "group"] +"""The two kinds of node a hierarchy holds.""" + + +def said(faults: Sequence[str]) -> str: + """`faults`, each a reason, as one clause: `a`, `a and b`, `a, b and c`.""" + if len(faults) <= 1: + return "".join(faults) + return f"{', '.join(faults[:-1])} and {faults[-1]}" + + +def name_faults(name: str) -> list[str]: + """What keeps `name` from being a node name, each said; none for a node name, the root's `""` among them.""" + faults: list[str] = [] + if "/" in name: + faults.append('holds "/"') + if name != "" and set(name) == {"."}: + faults.append("is periods alone") + if name.startswith("__"): + faults.append('starts with the reserved "__"') + if name == "zarr.json": + faults.append('is the reserved "zarr.json"') + return faults + + +def names_faults(names: list[str]) -> list[str]: + """What keeps `names`, the names a path holds between its `/`, from being node names below the root: the first that is not one, said, and how many more there are.""" + faulty = [name for name in names if name == "" or len(name_faults(name)) != 0] + if len(faulty) == 0: + return [] + first = faulty[0] + faults = ( + ['holds an empty name between two "/"'] + if first == "" + else [f"holds {shown(first)}, a name that {said(name_faults(first))}"] + ) + if len(faulty) > 1: + more = len(faulty) - 1 + faults.append( + f"holds {more} more name{'s' if more != 1 else ''} that {'are' if more != 1 else 'is'} not a node name" + ) + return faults + + +def path_faults(path: str) -> list[str]: + """What keeps `path` from being a node path, each said; none for a node path, the root's `/` among them.""" + if path == "/": + return [] + if not path.startswith("/"): + return ['does not start with "/"'] + faults: list[str] = [] + names = path[1:].split("/") + if path.endswith("/"): + faults.append('ends with "/"') + names = names[:-1] + return [*faults, *names_faults(names)] + + +def hierarchy_problems(nodes: Mapping[str, NodeType | None]) -> tuple[ValidationProblem, ...]: + """What keeps `nodes`, the node type of each node by its path, from being a Zarr v3 hierarchy: one problem per node it is about, at that node's path. + + "A Zarr hierarchy is a tree structure, where each node in the tree is + either a group or an array. Group nodes may have children but array + nodes may not" + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L177-L181). + So each path is a node path, and each node but the root has its parent + among them, a group: a node below an array is an `invalid_value` at its + path, and the nearest group missing above a node is a `missing_key` at + that group's path, once, saying how many more are missing above it, the + root's among them. A node of no type known, None, is taken as a group: + what makes its type unknown is its own problem, reported where it sits. + A mapping holds each path once, so no two siblings share a name. + """ + problems: list[ValidationProblem] = [] + well_formed: list[str] = [] + for path in nodes: + faults = path_faults(path) + if len(faults) == 0: + well_formed.append(path) + else: + message = f"expected a node path, got {shown(path)}, which {said(faults)}" + problems.append(ValidationProblem((path,), message, "invalid_value")) + tree = _Tree() + for path in well_formed: + tree.hold(path, nodes[path]) + reported: set[str] = set() + for path in well_formed: + if path == "/": + continue + holder, missing = tree.above(path) + if holder == "array": + # No node is below an array, however many groups are missing + # between them. + message = ( + f"expected a node below a group, got {shown(path)}, below the array " + f"{shown(tree.holder_of(path))}" + ) + problems.append(ValidationProblem((path,), message, "invalid_value")) + elif missing != 0: + nearest = path.rpartition("/")[0] or "/" + if nearest not in reported: + reported.add(nearest) + above = missing - 1 + more = ( + "" if above == 0 else f", and {above} group{'s' if above != 1 else ''} above it" + ) + message = f"missing the group holding {shown(path)}{more}" + problems.append(ValidationProblem((nearest,), message, "missing_key")) + return tuple(problems) + + +class _Branch: + """A name in the tree of held paths: whether a node is held there, its type, and the names below it.""" + + __slots__ = ("children", "held", "node_type") + + def __init__(self) -> None: + self.held = False + self.node_type: NodeType | None = None + self.children: dict[str, _Branch] = {} + + +class _Tree: + """The held paths as a tree of their names, so each path is walked once, in time proportional to its length. + + Walking up a path by its ancestors' strings costs the path's length + for each ancestor, which a hostile key of a million characters makes + a quadratic wait; here a path is split once and walked name by name. + """ + + __slots__ = ("root",) + + def __init__(self) -> None: + self.root = _Branch() + + def hold(self, path: str, node_type: NodeType | None) -> None: + """Mark the node at `path`, a node path, as held, of `node_type`.""" + branch = self.root + for name in _names(path): + branch = branch.children.setdefault(name, _Branch()) + branch.held = True + branch.node_type = node_type + + def above(self, path: str) -> tuple[NodeType | None, int]: + """Of the node at `path`, a node path below the root: the type of the nearest held node above it, `"group"` for one of no type known and None when none is held, and how many groups are missing between them.""" + names = _names(path) + branch: _Branch | None = self.root + holder: NodeType | None = None + held_depth = -1 + for depth, name in enumerate(names[:-1], start=0): + if branch is not None and branch.held: + holder, held_depth = branch.node_type or "group", depth + branch = None if branch is None else branch.children.get(name) + # The parent, at the last depth walked, may be held too. + if branch is not None and branch.held: + holder, held_depth = branch.node_type or "group", len(names) - 1 + return holder, len(names) - 1 - held_depth + + def holder_of(self, path: str) -> str: + """The path of the nearest held node above the node at `path`, a node path below the root that has one.""" + names = _names(path) + branch: _Branch | None = self.root + holder = "/" + for depth, name in enumerate(names[:-1]): + if branch is not None and branch.held: + holder = "/" + "/".join(names[:depth]) if depth != 0 else "/" + branch = None if branch is None else branch.children.get(name) + if branch is not None and branch.held: + holder = "/" + "/".join(names[:-1]) + return holder + + +def _names(path: str) -> list[str]: + """The names a node path holds between its `/`: none for the root's.""" + return [] if path == "/" else path[1:].split("/") + + +def _string_problems(value: object, what: str) -> tuple[ValidationProblem, ...]: + return (ValidationProblem((), f"expected {what}, got {shown(value)}", "invalid_type"),) + + +def validate_node_name_v3(value: object) -> tuple[ValidationProblem, ...]: + """Every reason `value` is not a v3 node name, said in one `invalid_value`, or an `invalid_type` for what is not a string.""" + if not isinstance(value, str): + return with_input(_string_problems(value, "a node name"), value) + faults = name_faults(value) + if len(faults) == 0: + return () + message = f"expected a node name, got {shown(value)}, which {said(faults)}" + return with_input((ValidationProblem((), message, "invalid_value"),), value) + + +def is_node_name_v3(value: object) -> TypeGuard[NodeName]: + """Whether `value` is a v3 node name `validate_node_name_v3` finds nothing wrong with.""" + return len(validate_node_name_v3(value)) == 0 + + +def parse_node_name_v3(value: object) -> NodeName: + """`value` as a `NodeName`, or `MetadataValidationError` with every reason it is not one.""" + problems = validate_node_name_v3(value) + if len(problems) != 0: + raise MetadataValidationError(problems) + return NodeName(cast("str", value)) + + +def validate_node_path_v3(value: object) -> tuple[ValidationProblem, ...]: + """Every reason `value` is not a v3 node path, said in one `invalid_value`, or an `invalid_type` for what is not a string.""" + if not isinstance(value, str): + return with_input(_string_problems(value, "a node path"), value) + faults = path_faults(value) + if len(faults) == 0: + return () + message = f"expected a node path, got {shown(value)}, which {said(faults)}" + return with_input((ValidationProblem((), message, "invalid_value"),), value) + + +def is_node_path_v3(value: object) -> TypeGuard[NodePath]: + """Whether `value` is a v3 node path `validate_node_path_v3` finds nothing wrong with.""" + return len(validate_node_path_v3(value)) == 0 + + +def parse_node_path_v3(value: object) -> NodePath: + """`value` as a `NodePath`, or `MetadataValidationError` with every reason it is not one.""" + problems = validate_node_path_v3(value) + if len(problems) != 0: + raise MetadataValidationError(problems) + return NodePath(cast("str", value)) + + +__all__ = [ + "NodeName", + "NodePath", + "NodeType", + "hierarchy_problems", + "is_node_name_v3", + "is_node_path_v3", + "name_faults", + "names_faults", + "parse_node_name_v3", + "parse_node_path_v3", + "path_faults", + "said", + "validate_node_name_v3", + "validate_node_path_v3", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/_pipeline.py b/packages/zarr-metadata/src/zarr_metadata/v3/_pipeline.py index a36017dc73..1b7857b5e6 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/_pipeline.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/_pipeline.py @@ -26,15 +26,24 @@ from __future__ import annotations import dataclasses -from collections.abc import Mapping from dataclasses import dataclass +from types import MappingProxyType from typing import TYPE_CHECKING, Any, Final, TypeGuard, cast -from zarr_metadata._json import ValidationProblem -from zarr_metadata.v3._definition import Chunk, CodecDefinition, CodecKind, Resolved, asked, ruled +from zarr_metadata._json import ValidationProblem, is_object, is_tuple, with_input +from zarr_metadata.v3._definition import ( + AcceptedField, + Chunk, + CodecDefinition, + CodecKind, + ResolvedField, + asked, + read_only, + ruled, +) if TYPE_CHECKING: - from collections.abc import Iterator, Sequence + from collections.abc import Iterator, Mapping, Sequence from zarr_metadata._typed_json import Loc from zarr_metadata.v3._definition import Nested, Problems @@ -49,7 +58,7 @@ def _no_stages() -> Mapping[str, tuple[Stage, ...]]: class Stage: """One codec of a pipeline, and the chunk it is handed.""" - codec: Resolved[CodecDefinition[Any]] + codec: ResolvedField[CodecDefinition[Any]] """The codec, as the scope read it.""" incoming: Chunk | None """The chunk it is handed. @@ -61,6 +70,14 @@ class Stage: inner: Mapping[str, tuple[Stage, ...]] = dataclasses.field(default_factory=_no_stages) """The pipelines it holds, by the member of its configuration that holds each: each codec with the chunk it is handed.""" + def __post_init__(self) -> None: + # Read-only, as everything a reading hands out is. + object.__setattr__(self, "inner", MappingProxyType(dict(self.inner))) + + def __reduce__(self) -> tuple[type[Stage], tuple[object, ...]]: + # A read-only view does not pickle: the stage pickles as what it was built from. + return Stage, (self.codec, self.incoming, dict(self.inner)) + _POSITIONS: Final[Mapping[CodecKind, int]] = { "array_array": 0, @@ -77,7 +94,7 @@ class Stage: def read_pipeline( - codecs: Sequence[Resolved[CodecDefinition[Any]]], chunk: Chunk, loc: Loc = () + codecs: Sequence[ResolvedField[CodecDefinition[Any]]], chunk: Chunk, loc: Loc = () ) -> tuple[tuple[Stage, ...], Problems]: """Each of `codecs`, codec fields a scope read, as a pipeline handed `chunk`, with the chunk it is handed, and what is wrong with them. @@ -95,7 +112,7 @@ def read_pipeline( # array -- past the array -> bytes codec, or a codec of unknown kind. handed: Chunk | None = chunk for index, codec in enumerate(codecs): - definition, configuration = codec.definition, codec.configuration + definition = codec.definition if definition is None: stages.append(Stage(codec, handed)) handed = None @@ -107,22 +124,25 @@ def read_pipeline( incoming = Chunk() if handed is None else handed at = (*loc, index, "configuration") inner = _no_stages() - if configuration is not None: - problems.extend(_chunk_problems(definition, configuration, codec.nested, incoming, at)) - inner, found = _inner_pipelines(definition, configuration, codec.nested, incoming, at) + if isinstance(codec, AcceptedField): + configuration, nested = codec.configuration, codec.nested + problems.extend(_chunk_problems(definition, configuration, nested, incoming, at)) + inner, found = _inner_pipelines(definition, configuration, nested, incoming, at) problems.extend(found) stages.append(Stage(codec, incoming, inner)) if definition.kind == "array_bytes": handed = None - elif configuration is None: + elif not isinstance(codec, AcceptedField): handed = Chunk() else: - handed = _handed_on(definition, configuration, codec.nested, incoming, at) - return tuple(stages), tuple(problems) + handed = _handed_on(definition, codec.configuration, codec.nested, incoming, at) + if len(problems) == 0: + return tuple(stages), () + return tuple(stages), with_input(problems, tuple(codec.json for codec in codecs), loc) def _order_problems( - codecs: Sequence[Resolved[CodecDefinition[Any]]], loc: Loc + codecs: Sequence[ResolvedField[CodecDefinition[Any]]], loc: Loc ) -> Iterator[ValidationProblem]: """Array -> array codecs, then one array -> bytes codec, then bytes -> bytes codecs. @@ -174,7 +194,9 @@ def _chunk_problems( at: Loc, ) -> Problems: """What `definition`'s chunk rules find in a codec handed `chunk`, located under `at`.""" - return ruled(definition, lambda: definition.chunk_rules(configuration, nested, chunk), at) + return ruled( + definition, lambda: definition.chunk_rules(read_only(configuration), nested, chunk), at + ) def _inner_pipelines( @@ -194,7 +216,7 @@ def _inner_pipelines( given = asked( definition, "pipelines", - lambda: cast("object", definition.pipelines(configuration, nested, chunk)), + lambda: cast("object", definition.pipelines(read_only(configuration), nested, chunk)), at, ) if not _is_pipelines(given): @@ -213,35 +235,23 @@ def _inner_pipelines( def _is_pipelines(value: object) -> TypeGuard[Mapping[str, Chunk]]: """Whether `value` maps members of a configuration to chunks.""" - return isinstance(value, Mapping) and all( - isinstance(member, str) and isinstance(chunk, Chunk) - for member, chunk in cast("Mapping[object, object]", value).items() + return is_object(value) and all( + isinstance(member, str) and isinstance(chunk, Chunk) for member, chunk in value.items() ) def _held( definition: CodecDefinition[Any], configuration: Mapping[str, Any], nested: Nested, member: str -) -> tuple[Resolved[CodecDefinition[Any]], ...]: +) -> tuple[ResolvedField[CodecDefinition[Any]], ...]: """The codecs `member` of the configuration holds, as the scope read them. - A member holding anything but a list of fields, or fields a definition - of another kind read, is a `TypeError`: a fault in the definition that - names it. Fields nothing in scope claims tell nothing of their kind, - and are read as codecs nothing claims. + A member holding anything but a list of fields read as codecs is a + `TypeError`: a fault in the definition that names it. """ entries: object = configuration.get(member) - places = ( - [(member, index) for index in range(len(cast("tuple[object, ...]", entries)))] - if isinstance(entries, tuple) - else None - ) + places = [(member, index) for index in range(len(entries))] if is_tuple(entries) else None if places is None or not all( - place in nested - and ( - nested[place].definition is None - or isinstance(nested[place].definition, CodecDefinition) - ) - for place in places + place in nested and nested[place].read_as is CodecDefinition for place in places ): msg = f"{definition.name!r}: its pipelines name {member!r}, which holds no list of codecs" raise TypeError(msg) @@ -264,7 +274,7 @@ def _handed_on( given = asked( definition, "transition", - lambda: cast("object", definition.transition(configuration, nested, chunk)), + lambda: cast("object", definition.transition(read_only(configuration), nested, chunk)), at, ) if not isinstance(given, Chunk): diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/_registry.py b/packages/zarr-metadata/src/zarr_metadata/v3/_registry.py index e5beff8c64..38c9dfc012 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/_registry.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/_registry.py @@ -11,7 +11,7 @@ uses nothing an implementation could refuse for being optional. `CORE_AND_EXTENSIONS` adds what `zarr-extensions` registers and this package defines. A name in neither is not refused -- that is what keeps -the format open -- it is read as `out_of_scope` and left unjudged. +the format open -- it is read as `UnclaimedField` and left unjudged. """ from __future__ import annotations @@ -19,9 +19,12 @@ from collections.abc import Mapping from dataclasses import dataclass from types import MappingProxyType -from typing import TYPE_CHECKING, Any, Final, cast +from typing import TYPE_CHECKING, Any, Final, Generic, Literal, TypeAlias, TypeGuard, cast -from zarr_metadata.v3._definition import KINDS, Definition, as_kind, kind_of, spelled +from typing_extensions import TypeVar + +from zarr_metadata.v3._definition import Definition, as_kind, format_of, kind_of, spelled +from zarr_metadata.v3._scope import Conflict, ScopeConflictError, disagreements_of from zarr_metadata.v3.chunk_grid.rectilinear import RECTILINEAR_CHUNK_GRID from zarr_metadata.v3.chunk_grid.regular import REGULAR_CHUNK_GRID from zarr_metadata.v3.chunk_key_encoding.default import DEFAULT_CHUNK_KEY_ENCODING @@ -57,49 +60,75 @@ from zarr_metadata.v3.data_type.uint64 import UINT64_DATA_TYPE if TYPE_CHECKING: + from collections.abc import Callable + from zarr_metadata.v3._definition import D + from zarr_metadata.v3._scope import Claims, Disagreements Tables = Mapping[type[Definition[Any]], Mapping[str, Definition[Any]]] """By kind, then by the name each definition is filed under.""" +F = TypeVar("F", bound=Literal[2, 3], default=Any) +"""The Zarr format a scope reads documents of: 2 or 3, and any when a scope is built from definitions rather than from a format's own.""" -@dataclass(frozen=True, slots=True) -class Context: - """The definitions in scope while metadata is read. + +@dataclass(frozen=True, slots=True, eq=False) +class Context(Generic[F]): + """The definitions in scope while metadata is read, of one Zarr format. A value with no reading of its own: `resolve` reads a field in it, and `claimant` is the one question it answers, which definition a name belongs to. Built from definitions with `Context.of`, extended - with more by `extended_with`. + with more by `extended_with`; two scopes are equal when they file the + same definitions, and equal scopes hash alike. Every kind a scope + files declares one format, which is the scope's: a v3 reader refuses + a v2 scope, and the other way round, while a scope that files nothing + is of no format and reads in either, claiming nothing. """ tables: Tables + def __post_init__(self) -> None: + # What `of` builds, the constructor refuses to build otherwise: each + # key a kind, and every kind of one format. + formats = sorted({format_of(kind) for kind in self.tables}) + if len(formats) > 1: + msg = ( + "a scope reads documents of one Zarr format; these definitions are of kinds of " + f"formats {' and '.join(f'v{found}' for found in formats)}" + ) + raise TypeError(msg) + + @property + def format(self) -> Literal[2, 3] | None: + """The Zarr format the definitions in scope read, 2 or 3; None for a scope that files nothing.""" + return next((format_of(kind) for kind in self.tables), None) + @classmethod - def of(cls, *definitions: Definition[Any]) -> Context: + def of(cls, *definitions: Definition[Any]) -> Context[Any]: """A scope of exactly these definitions; a later one takes a name over from an earlier. `TypeError` for a definition of no kind, which no position in a - document could hold. + document could hold, and for definitions of kinds of two formats, + which no document holds together. """ - tables: dict[type[Definition[Any]], dict[str, Definition[Any]]] = { - kind: {} for kind in KINDS - } + tables: dict[type[Definition[Any]], dict[str, Definition[Any]]] = {} for definition in definitions: kind = kind_of(definition) if kind is None: msg = ( f"{definition.name!r} is a definition of no kind; build it as a " "CodecDefinition, DataTypeDefinition, ChunkGridDefinition, " - "ChunkKeyEncodingDefinition or StorageTransformerDefinition" + "ChunkKeyEncodingDefinition or StorageTransformerDefinition, or as a " + "kind of your own" ) raise TypeError(msg) - tables[kind][definition.name] = definition + tables.setdefault(kind, {})[definition.name] = definition return cls( MappingProxyType({kind: MappingProxyType(table) for kind, table in tables.items()}) ) - def extended_with(self, *definitions: Definition[Any]) -> Context: + def extended_with(self, *definitions: Definition[Any]) -> Context[F]: """This scope, plus definitions of your own. A name already filed under the same kind is taken over by what is @@ -113,6 +142,70 @@ def definitions(self) -> tuple[Definition[Any], ...]: """Every definition in scope, kind by kind.""" return tuple(entry for table in self.tables.values() for entry in table.values()) + def __repr__(self) -> str: + # Short, as a default argument shows it: in full, a scope's repr is + # every definition's, and `help` of a validator runs to pages. + return f"Context(<{len(self.definitions())} definitions>)" + + def __reduce__(self) -> tuple[Callable[..., Context], tuple[Definition[Any], ...]]: + # A scope is its definitions, so it pickles as them, and goes to + # another process with the documents it is to read there. + return (Context.of, self.definitions()) + + def __copy__(self) -> Context: + return self + + def __deepcopy__(self, memo: dict[int, object]) -> Context: + # A scope never changes, so a copy of it is itself. + return self + + def __eq__(self, other: object) -> bool: + if not isinstance(other, Context): + return NotImplemented + return self._filed() == other._filed() + + def __hash__(self) -> int: + return hash(self._filed()) + + def _filed(self) -> frozenset[tuple[type[Definition[Any]], str, Definition[Any]]]: + """Every definition in scope with the kind and name it is filed under: what two scopes are compared by.""" + return frozenset( + (kind, name, definition) + for kind, table in self.tables.items() + for name, definition in table.items() + ) + + def disagreements(self, claims: Claims) -> Disagreements: + """Where this scope reads `claims`, a reading's, otherwise: what it would gain, and what it conflicts with, as `Disagreements` says. + + A claim is keyed by the name its definition is filed under -- raw + bits under `r*` -- so it is looked up as filed, not as a document + writes it. + """ + return disagreements_of(lambda kind, name: self.tables.get(kind, {}).get(name), claims) + + @classmethod + def joined(cls, *contexts: Context[F]) -> Context[F]: + """The least scope that files everything each of `contexts` files: their join. + + `ScopeConflictError` when two of them file different definitions + under one name of one kind; `extended_with` is for taking a name + over on purpose. + """ + filed: dict[tuple[type[Definition[Any]], str], Definition[Any]] = {} + conflicts: list[Conflict] = [] + for context in contexts: + for kind, table in context.tables.items(): + for name, definition in table.items(): + held = filed.get((kind, name)) + if held is not None and held != definition: + conflicts.append(Conflict((kind, name), held, definition)) + continue + filed[kind, name] = definition + if len(conflicts) != 0: + raise ScopeConflictError(conflicts) + return cls.of(*filed.values()) + def claimant(self, kind: type[D], name: str) -> D | None: """The definition of `kind` in scope that reads `name`, a name a document writes; None if none does. @@ -126,6 +219,30 @@ def claimant(self, kind: type[D], name: str) -> D | None: return cast("D | None", self.tables.get(asked, {}).get(filed)) +ZarrV3Context: TypeAlias = Context[Literal[3]] +"""A scope that reads Zarr v3 documents: what every v3 reader takes.""" +ZarrV2Context: TypeAlias = Context[Literal[2]] +"""A scope that reads Zarr v2 documents: what every v2 reader takes.""" + + +def is_scope(value: object) -> TypeGuard[Context[Any]]: + """Whether `value` is a scope, of whatever format.""" + return isinstance(value, Context) + + +def scoped(context: Context[F] | None, default: Context[F]) -> Context[F]: + """The scope a reader of `default`'s format reads in: `context`, or `default` when none is given; `TypeError` for a scope of the other format.""" + if context is None: + return default + if context.format is not None and context.format != default.format: + msg = ( + f"a v{default.format} document is read in a scope of that format or of none, " + f"got a v{context.format} scope" + ) + raise TypeError(msg) + return context + + _CORE: Final[tuple[Definition[Any], ...]] = ( BLOSC_CODEC, BYTES_CODEC, @@ -167,10 +284,19 @@ def claimant(self, kind: type[D], name: str) -> D | None: ) """What `zarr-extensions` registers and this package defines.""" -CORE: Final = Context.of(*_CORE) +CORE: Final[ZarrV3Context] = Context.of(*_CORE) """Only what the Zarr v3 specification defines.""" -CORE_AND_EXTENSIONS: Final = Context.of(*_CORE, *_EXTENSIONS) +CORE_AND_EXTENSIONS: Final[ZarrV3Context] = Context.of(*_CORE, *_EXTENSIONS) """What the specification defines, plus what `zarr-extensions` registers.""" -__all__ = ["CORE", "CORE_AND_EXTENSIONS", "Context", "Tables"] +__all__ = [ + "CORE", + "CORE_AND_EXTENSIONS", + "Context", + "Tables", + "ZarrV2Context", + "ZarrV3Context", + "is_scope", + "scoped", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/_scope.py b/packages/zarr-metadata/src/zarr_metadata/v3/_scope.py new file mode 100644 index 0000000000..5aa75c8e02 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v3/_scope.py @@ -0,0 +1,230 @@ +"""The algebra of scopes: what a reading claims of each name, the order one reading refines another in, and where two scopes disagree. + +A model is a document read in a scope. What the document means depends +only on what the scope said of the names it writes -- its *claims* -- +not on everything the scope holds. Two readings of one document are +ordered by information: a name nothing claimed, read by a definition, +gains meaning and loses none; a name read by one definition and then +another is a conflict. `Context.disagreements` and `Context.joined` are +built on these. +""" + +from __future__ import annotations + +from collections.abc import Mapping +from dataclasses import dataclass +from typing import TYPE_CHECKING, Any, Literal, TypeAlias, cast + +from zarr_metadata.v3._definition import ( + AcceptedField, + Definition, + RefusedField, + UnclaimedField, + as_kind, + field_key, + fields_of, + format_of, + own_key, + resolve, + spelled, +) + +if TYPE_CHECKING: + from collections.abc import Callable, Iterable, Sequence + + from zarr_metadata._typed_json import Loc + from zarr_metadata.v3._definition import D, ResolvedField + +ClaimKey: TypeAlias = tuple[type[Definition[Any]], str] +"""A kind and the name a definition is filed under: what a scope answers `claimant` for.""" + +Claims: TypeAlias = Mapping[ClaimKey, Definition[Any] | None] +"""What a reading claims of each name a document writes: the definition that read it, or None where nothing claimed it.""" + + +@dataclass(frozen=True, slots=True) +class Conflict: + """One place two readings of a name disagree: the key, what one claimed, what the other found, and where in a document when known.""" + + key: ClaimKey + claimed: Definition[Any] | None + found: Definition[Any] | None + loc: Loc | None = None + written: str | None = None + """The name the document writes where the conflict was found, `r16` for a key of `r*`; None when no document is at hand.""" + + def __str__(self) -> str: + kind, filed = self.key + name = filed if self.written is None else self.written + where = "" if self.loc is None else f" at {self.loc!r}" + return ( + f"{kind_name(kind)} {name!r}{where}: claimed {definition_said(self.claimed)}, " + f"found {definition_said(self.found)}" + ) + + +def definition_said(definition: Definition[Any] | None) -> str: + """A definition as a message tells it from another of the same name: by the TypedDict its configuration is; "no definition" for None.""" + if definition is None: + return "no definition" + return f"{definition!r} of {definition.configuration.__qualname__}" + + +def kind_name(kind: type[Definition[Any]]) -> str: + """A kind of definition as a message names it: its label, "codec" for `CodecDefinition`.""" + return kind.label + + +class ScopeConflictError(ValueError): + """Raised where two scopes, or a scope and a reading, give one name two meanings, or one would lose a meaning the other has. + + Carries every conflict in `.conflicts`, as `MetadataValidationError` + carries every problem. + """ + + def __init__(self, conflicts: Sequence[Conflict]) -> None: + self.conflicts: tuple[Conflict, ...] = tuple(conflicts) + super().__init__("; ".join(str(conflict) for conflict in self.conflicts)) + + def __reduce__( + self, + ) -> tuple[type[ScopeConflictError], tuple[tuple[Conflict, ...]], dict[str, object]]: + # Pickled and copied as it was raised, its notes and attributes kept: + # an exception's default reduce calls the constructor with its + # message, not its conflicts. + return type(self), (self.conflicts,), dict(self.__dict__) + + +def claim_key(field: ResolvedField[Any]) -> ClaimKey | None: + """The key `field` is claimed under: its kind and the name its definition is filed under, `r*` for `r16`; None for a field that names nothing.""" + if field.name is None: + return None + filed, _ = spelled(field.read_as, field.name) + return None if filed is None else (field.read_as, filed) + + +def claims_of( + fields: Iterable[tuple[Loc, ResolvedField[Any]]], +) -> dict[ClaimKey, Definition[Any] | None]: + """What `fields`, each with where it sits, claim of each name: the definition that read it, None where nothing claimed it. + + `fields` are one reading's, as `fields_of` or a reading's `fields()` + gives them. A name claimed two ways among them is a + `ScopeConflictError`: no one scope read them. + """ + claims: dict[ClaimKey, Definition[Any] | None] = {} + conflicts: list[Conflict] = [] + for loc, field in fields: + key = claim_key(field) + if key is None: + continue + definition = field.definition + if key in claims and claims[key] != definition: + conflicts.append(Conflict(key, claims[key], definition, loc)) + continue + claims[key] = definition + if len(conflicts) != 0: + raise ScopeConflictError(conflicts) + return claims + + +def refines(field: ResolvedField[Any], other: ResolvedField[Any]) -> bool: + """Whether `field` holds everything `other` holds: reads the same where both read, and reads what `other` left unclaimed. + + The order one reading of a document refines another in. A name nothing + claimed, read by a definition, is a gain when what the unclaimed field + wrote, read by that definition and those of the fields the read field + holds, is the read field: two spellings of one configuration are one + gain, as they are one field to `==`, so the order is transitive + through equality. The reverse is a loss; one name read by two + definitions is a conflict; a refused field refines itself alone. Two + fields that refine each other are equal. + """ + if isinstance(field, RefusedField) or isinstance(other, RefusedField): + return field == other + if isinstance(other, UnclaimedField): + if isinstance(field, UnclaimedField): + return field_key(field) == field_key(other) + return claim_key(field) == claim_key(other) and _gained(field, other) + if isinstance(field, UnclaimedField): + return False + if field.definition != other.definition or own_key(field) != own_key(other): + return False + if set(field.nested) != set(other.nested): + return False + return all(refines(field.nested[loc], other.nested[loc]) for loc in field.nested) + + +def _gained(field: AcceptedField[Any], other: UnclaimedField) -> bool: + """Whether `other`, read in the scope `field`'s own claims make, is `field`: what a gain is.""" + again, _ = resolve(other.json, field.read_as, _Claimed.of(field)) + return again == field + + +@dataclass(frozen=True, slots=True) +class _Claimed: + """The scope a field's own claims make: its definition and those of the fields it holds, by kind and filed name; what a gain re-reads in.""" + + filed: Mapping[ClaimKey, Definition[Any]] + format: Literal[2, 3] | None + + @classmethod + def of(cls, field: AcceptedField[Any]) -> _Claimed: + claimed = claims_of(fields_of(field)) + return cls( + {key: definition for key, definition in claimed.items() if definition is not None}, + format_of(field.read_as), + ) + + def claimant(self, kind: type[D], name: str) -> D | None: + """The definition of `kind` among the claims that reads `name`; None if none does, as `Context.claimant` answers.""" + asked = as_kind(kind) + filed, _ = spelled(asked, name) + if filed is None: + return None + return cast("D | None", self.filed.get((asked, filed))) + + +@dataclass(frozen=True, slots=True) +class Disagreements: + """Where a scope reads a reading's claims otherwise: the names it would gain a meaning for, and those it conflicts with, a lost meaning among them.""" + + gains: tuple[ClaimKey, ...] + conflicts: tuple[Conflict, ...] + + @property + def agrees(self) -> bool: + """Whether the scope reads every claim identically.""" + return len(self.gains) == 0 and len(self.conflicts) == 0 + + +def disagreements_of( + claimant: Callable[[type[Definition[Any]], str], Definition[Any] | None], claims: Claims +) -> Disagreements: + """`Disagreements` between what `claimant`, asked by kind and filed name as a scope's tables answer, gives for each key and what `claims` records.""" + gains: list[ClaimKey] = [] + conflicts: list[Conflict] = [] + for key, claimed in claims.items(): + kind, name = key + found = claimant(kind, name) + if found == claimed: + continue + if claimed is None: + gains.append(key) + else: + conflicts.append(Conflict(key, claimed, found)) + return Disagreements(tuple(gains), tuple(conflicts)) + + +__all__ = [ + "ClaimKey", + "Claims", + "Conflict", + "Disagreements", + "ScopeConflictError", + "claim_key", + "claims_of", + "disagreements_of", + "kind_name", + "refines", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/array.py b/packages/zarr-metadata/src/zarr_metadata/v3/array.py index 31a5f6b755..79328d4eb7 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/array.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/array.py @@ -1,12 +1,19 @@ """Zarr v3 array metadata types.""" from collections.abc import Mapping -from typing import Final, Literal, NotRequired, TypeAlias +from typing import Annotated, Final, Literal, NotRequired, TypeAlias +from annotated_types import Ge from typing_extensions import TypedDict from zarr_metadata._common import JSONValue -from zarr_metadata.v3._common import ZarrV3MetadataFieldJSON +from zarr_metadata.v3._common import ( + ChunkGridField, + ChunkKeyEncodingField, + CodecField, + DataTypeField, + StorageTransformerField, +) ZarrV3ExtensionField: TypeAlias = JSONValue """The JSON value of an unknown top-level v3 metadata field. @@ -21,21 +28,24 @@ class ZarrV3ArrayMetadataJSON(TypedDict, extra_items=ZarrV3ExtensionField): """ Zarr v3 array metadata document (the `zarr.json` content for an array). - Extra keys may contain arbitrary JSON values. + Extra keys may contain arbitrary JSON values. Each extension point is + annotated with the field alias of its kind -- `data_type` a + `DataTypeField`, each of `codecs` a `CodecField` -- which is the JSON + a metadata field is, and says what a scope reads it as. See https://zarr-specs.readthedocs.io/en/latest/v3/core/index.html#array-metadata """ zarr_format: Literal[3] node_type: Literal["array"] - data_type: ZarrV3MetadataFieldJSON - shape: tuple[int, ...] - chunk_grid: ZarrV3MetadataFieldJSON - chunk_key_encoding: ZarrV3MetadataFieldJSON + data_type: DataTypeField + shape: tuple[Annotated[int, Ge(0)], ...] + chunk_grid: ChunkGridField + chunk_key_encoding: ChunkKeyEncodingField fill_value: JSONValue - codecs: tuple[ZarrV3MetadataFieldJSON, ...] + codecs: tuple[CodecField, ...] attributes: NotRequired[Mapping[str, JSONValue]] - storage_transformers: NotRequired[tuple[ZarrV3MetadataFieldJSON, ...]] + storage_transformers: NotRequired[tuple[StorageTransformerField, ...]] dimension_names: NotRequired[tuple[str | None, ...]] @@ -64,14 +74,14 @@ class ZarrV3ArrayMetadataJSONPartial(TypedDict, total=False, extra_items=ZarrV3E zarr_format: Literal[3] node_type: Literal["array"] - data_type: ZarrV3MetadataFieldJSON - shape: tuple[int, ...] - chunk_grid: ZarrV3MetadataFieldJSON - chunk_key_encoding: ZarrV3MetadataFieldJSON + data_type: DataTypeField + shape: tuple[Annotated[int, Ge(0)], ...] + chunk_grid: ChunkGridField + chunk_key_encoding: ChunkKeyEncodingField fill_value: JSONValue - codecs: tuple[ZarrV3MetadataFieldJSON, ...] + codecs: tuple[CodecField, ...] attributes: NotRequired[Mapping[str, JSONValue]] - storage_transformers: NotRequired[tuple[ZarrV3MetadataFieldJSON, ...]] + storage_transformers: NotRequired[tuple[StorageTransformerField, ...]] dimension_names: NotRequired[tuple[str | None, ...]] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/chunk_grid/rectilinear.py b/packages/zarr-metadata/src/zarr_metadata/v3/chunk_grid/rectilinear.py index a8191727a8..157b682e42 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/chunk_grid/rectilinear.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/chunk_grid/rectilinear.py @@ -5,12 +5,12 @@ """ from collections.abc import Iterator -from typing import Final, Literal, NotRequired +from typing import Annotated, Final, Literal, NotRequired +from annotated_types import Ge from typing_extensions import TypedDict from zarr_metadata._json import ValidationProblem -from zarr_metadata._typed_json import Loc from zarr_metadata.v3._definition import ChunkGridDefinition, Lengths, Nested RECTILINEAR_CHUNK_GRID_NAME: Final = "rectilinear" @@ -19,12 +19,14 @@ RectilinearChunkGridName = Literal["rectilinear"] """Literal type of the `name` field of the rectilinear chunk grid.""" -RectilinearDimSpec = int | tuple[int | tuple[int, int], ...] +_Positive = Annotated[int, Ge(1)] + +RectilinearDimSpec = _Positive | tuple[_Positive | tuple[_Positive, _Positive], ...] """JSON shape for one dimension's rectilinear spec. Either a bare integer (uniform shorthand for a regular dimension within a rectilinear grid), or a tuple of integers and/or `[value, count]` RLE -pairs. +pairs. Every extent, and every run's length and count, is at least 1. """ @@ -53,29 +55,6 @@ class RectilinearChunkGridObject(TypedDict, closed=True): """ -def _not_positive(loc: Loc, value: int) -> ValidationProblem: - return ValidationProblem(loc, f"expected an integer >= 1, got {value}", "invalid_value") - - -def _rules( - configuration: RectilinearChunkGridConfiguration, nested: Nested -) -> Iterator[ValidationProblem]: - """Every extent, and every run's length and count, is at least 1.""" - for axis, spec in enumerate(configuration["chunk_shapes"]): - if isinstance(spec, int): - if spec < 1: - yield _not_positive(("chunk_shapes", axis), spec) - continue - for index, entry in enumerate(spec): - if isinstance(entry, int): - if entry < 1: - yield _not_positive(("chunk_shapes", axis, index), entry) - continue - for position, value in enumerate(entry): - if value < 1: - yield _not_positive(("chunk_shapes", axis, index, position), value) - - def canonical_dim_spec(spec: RectilinearDimSpec) -> RectilinearDimSpec: """One dimension's chunk sizes in their simplest equivalent form. @@ -178,7 +157,6 @@ def _chunk_lengths( RECTILINEAR_CHUNK_GRID: Final = ChunkGridDefinition( name=RECTILINEAR_CHUNK_GRID_NAME, configuration=RectilinearChunkGridConfiguration, - rules=_rules, canonical=_canonical, shape_rules=_shape_rules, chunk_lengths=_chunk_lengths, diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/chunk_grid/regular.py b/packages/zarr-metadata/src/zarr_metadata/v3/chunk_grid/regular.py index 444ec11f4a..299603475d 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/chunk_grid/regular.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/chunk_grid/regular.py @@ -5,8 +5,9 @@ """ from collections.abc import Iterator -from typing import Final, Literal, NotRequired +from typing import Annotated, Final, Literal, NotRequired +from annotated_types import Ge from typing_extensions import TypedDict from zarr_metadata._json import ValidationProblem @@ -20,9 +21,18 @@ class RegularChunkGridConfiguration(TypedDict, closed=True): - """Configuration for the regular chunk grid.""" + """Configuration for the regular chunk grid. + + Every chunk length is at least 1, along a dimension of length 0 too: + "Chunk sizes must be greater than zero" + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-grids/regular-grid/index.rst#L40). + The core spec's "The chunk shape elements are non-zero when the + corresponding dimensions of the arrays have non-zero length" + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L284-L285) + says less of an empty dimension, and allows nothing the grid does not. + """ - chunk_shape: tuple[int, ...] + chunk_shape: tuple[Annotated[int, Ge(1)], ...] class RegularChunkGridObject(TypedDict, closed=True): @@ -43,35 +53,14 @@ class RegularChunkGridObject(TypedDict, closed=True): """ -def _rules( - configuration: RegularChunkGridConfiguration, nested: Nested -) -> Iterator[ValidationProblem]: - """No chunk extent is negative. - - "The chunk shape elements are non-zero when the corresponding - dimensions of the arrays have non-zero length": an extent of 0 is - right for a dimension of length 0, which zarr-python 3.0 and 3.1 - wrote, and which a grid alone cannot tell from one that is not; the - shape rules can, beside the array's shape. - """ - for index, extent in enumerate(configuration["chunk_shape"]): - if extent < 0: - yield ValidationProblem( - ("chunk_shape", index), f"expected an integer >= 0, got {extent}", "invalid_value" - ) - - def _shape_rules( configuration: RegularChunkGridConfiguration, nested: Nested, shape: tuple[int, ...] ) -> Iterator[ValidationProblem]: - """A chunk length for each of the array's dimensions, and 0 only for a dimension of length 0. + """A chunk length for each of the array's dimensions. "The dimensionality of the grid is the same as the dimensionality of the array" - (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-grids/regular-grid/index.rst#L29-L31), - and "The chunk shape elements are non-zero when the corresponding - dimensions of the arrays have non-zero length" - (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L284-L285). + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-grids/regular-grid/index.rst#L29-L31). """ chunk_shape = configuration["chunk_shape"] if len(chunk_shape) != len(shape): @@ -80,14 +69,6 @@ def _shape_rules( f"expected one chunk length per dimension of shape, got {len(chunk_shape)}", "invalid_value", ) - return - for axis, (length, extent) in enumerate(zip(chunk_shape, shape, strict=True)): - if length == 0 and extent != 0: - yield ValidationProblem( - ("chunk_shape", axis), - f"expected a chunk length >= 1 for a dimension of length {extent}, got 0", - "invalid_value", - ) def _chunk_lengths( @@ -107,7 +88,6 @@ def _chunk_lengths( REGULAR_CHUNK_GRID: Final = ChunkGridDefinition( name=REGULAR_CHUNK_GRID_NAME, configuration=RegularChunkGridConfiguration, - rules=_rules, shape_rules=_shape_rules, chunk_lengths=_chunk_lengths, ) diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/chunk_key_encoding/default.py b/packages/zarr-metadata/src/zarr_metadata/v3/chunk_key_encoding/default.py index c0e49ee543..cc6ab57062 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/chunk_key_encoding/default.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/chunk_key_encoding/default.py @@ -7,7 +7,7 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/core/index.html#chunk-key-encoding """ -from typing import Final, Literal, NotRequired +from typing import Final, Literal, NotRequired, cast from typing_extensions import TypedDict @@ -54,8 +54,25 @@ class DefaultChunkKeyEncodingObject(TypedDict, closed=True): so the short-hand-name form is permitted in addition to the object form. """ + +def _canonical( + configuration: DefaultChunkKeyEncodingConfiguration, +) -> DefaultChunkKeyEncodingConfiguration: + """Without a `separator` of `/`, which is what an absent one means. + + "If not specified, `separator` defaults to `/`" + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-key-encodings/default/index.rst#L27-L29), + so the two spellings are one encoding. + """ + if configuration.get("separator") != "/": + return configuration + return cast("DefaultChunkKeyEncodingConfiguration", {}) + + DEFAULT_CHUNK_KEY_ENCODING: Final = ChunkKeyEncodingDefinition( - name=DEFAULT_CHUNK_KEY_ENCODING_NAME, configuration=DefaultChunkKeyEncodingConfiguration + name=DEFAULT_CHUNK_KEY_ENCODING_NAME, + configuration=DefaultChunkKeyEncodingConfiguration, + canonical=_canonical, ) """The `default` chunk key encoding; its `separator` is typed, so it has no rule of its own.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/chunk_key_encoding/v2.py b/packages/zarr-metadata/src/zarr_metadata/v3/chunk_key_encoding/v2.py index 7c178a599e..4d9bc284c4 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/chunk_key_encoding/v2.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/chunk_key_encoding/v2.py @@ -13,7 +13,7 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/core/index.html#chunk-key-encoding """ -from typing import Final, Literal, NotRequired +from typing import Final, Literal, NotRequired, cast from typing_extensions import TypedDict @@ -60,8 +60,23 @@ class V2ChunkKeyEncodingObject(TypedDict, closed=True): so the short-hand-name form is permitted in addition to the object form. """ + +def _canonical(configuration: V2ChunkKeyEncodingConfiguration) -> V2ChunkKeyEncodingConfiguration: + """Without a `separator` of `.`, which is what an absent one means. + + "If not specified, `separator` defaults to `.`" + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/chunk-key-encodings/v2/index.rst#L27-L29), + so the two spellings are one encoding. + """ + if configuration.get("separator") != ".": + return configuration + return cast("V2ChunkKeyEncodingConfiguration", {}) + + V2_CHUNK_KEY_ENCODING: Final = ChunkKeyEncodingDefinition( - name=V2_CHUNK_KEY_ENCODING_NAME, configuration=V2ChunkKeyEncodingConfiguration + name=V2_CHUNK_KEY_ENCODING_NAME, + configuration=V2ChunkKeyEncodingConfiguration, + canonical=_canonical, ) """The `v2` chunk key encoding; its `separator` is typed, so it has no rule of its own.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/codec/_arithmetic.py b/packages/zarr-metadata/src/zarr_metadata/v3/codec/_arithmetic.py index 47dd9d27f1..4cf96bb904 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/codec/_arithmetic.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/codec/_arithmetic.py @@ -17,9 +17,9 @@ which both codecs do, while neither lists them. """ -from typing import Any, Final +from typing import Final -from zarr_metadata.v3._definition import RAW_BYTES_NAME, Resolved +from zarr_metadata.v3._definition import RAW_BYTES_NAME from zarr_metadata.v3.data_type.bool import BOOL_DATA_TYPE_NAME from zarr_metadata.v3.data_type.bytes import BYTES_DATA_TYPE_NAME from zarr_metadata.v3.data_type.complex64 import COMPLEX64_DATA_TYPE_NAME @@ -67,11 +67,4 @@ """ -def read_name(data_type: Resolved[Any] | None) -> str | None: - """The name the definition that read `data_type` is filed under; None when no definition read it.""" - if data_type is None or data_type.resolution != "read" or data_type.definition is None: - return None - return data_type.definition.name - - -__all__ = ["COMPLEX", "FLOATING_POINT", "NOT_NUMBERS", "read_name"] +__all__ = ["COMPLEX", "FLOATING_POINT", "NOT_NUMBERS"] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/codec/blosc.py b/packages/zarr-metadata/src/zarr_metadata/v3/codec/blosc.py index fb162c537d..c5aed10ced 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/codec/blosc.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/codec/blosc.py @@ -5,8 +5,9 @@ """ from collections.abc import Iterator -from typing import Final, Literal, NotRequired, cast +from typing import Annotated, Final, Literal, NotRequired, cast +from annotated_types import Ge, Interval from typing_extensions import TypedDict from zarr_metadata._json import ValidationProblem @@ -35,9 +36,9 @@ class BloscCodecConfiguration(TypedDict, closed=True): """Configuration for the Zarr v3 `blosc` codec.""" cname: BloscCName - clevel: int + clevel: Annotated[int, Interval(ge=0, le=9)] shuffle: BloscShuffle - blocksize: int + blocksize: Annotated[int, Ge(0)] typesize: NotRequired[int] @@ -69,21 +70,13 @@ class BloscCodecObject(TypedDict, closed=True): def _rules(configuration: BloscCodecConfiguration, nested: Nested) -> Iterator[ValidationProblem]: - """Bounds on `clevel` and `blocksize`; `typesize` against `shuffle`. + """`typesize` against `shuffle`. Under `noshuffle` the spec says of `typesize` that "the value is ignored", and the canonical form drops it; under either shuffle it is - required, and positive. + required, and positive. Whether it is required is not the type's to + say, so neither is its bound. """ - clevel, blocksize = configuration["clevel"], configuration["blocksize"] - if not 0 <= clevel <= 9: - yield ValidationProblem( - ("clevel",), f"expected an integer in [0, 9], got {clevel}", "invalid_value" - ) - if blocksize < 0: - yield ValidationProblem( - ("blocksize",), f"expected an integer >= 0, got {blocksize}", "invalid_value" - ) shuffle = configuration["shuffle"] if shuffle != BLOSC_NO_SHUFFLE: typesize = configuration.get("typesize") @@ -93,7 +86,10 @@ def _rules(configuration: BloscCodecConfiguration, nested: Nested) -> Iterator[V ) elif typesize < 1: yield ValidationProblem( - ("typesize",), f"expected a positive integer, got {typesize}", "invalid_value" + ("typesize",), + f"expected a positive integer, got {typesize}", + "invalid_value", + ctx={"ge": 1}, ) diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/codec/bytes.py b/packages/zarr-metadata/src/zarr_metadata/v3/codec/bytes.py index c98f1d712d..cbd3e5cdf3 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/codec/bytes.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/codec/bytes.py @@ -9,12 +9,11 @@ from typing_extensions import TypedDict -from zarr_metadata._json import ValidationProblem +from zarr_metadata._json import ValidationProblem, shown from zarr_metadata.v3._definition import ( Chunk, CodecDefinition, Nested, - named_configuration, storage_of, ) @@ -91,7 +90,7 @@ def _chunk_rules( if chunk.data_type is None: return storage = storage_of(chunk.data_type) - written, _, _ = named_configuration(chunk.data_type.json) + written = chunk.data_type.name if storage == "multi_byte" and "endian" not in configuration: yield ValidationProblem( ("endian",), @@ -101,7 +100,7 @@ def _chunk_rules( elif storage == "variable_length": yield ValidationProblem( (), - f"expected a data type of fixed size, got {written!r}, whose values vary in size", + f"expected a data type of fixed size, got {shown(written)}, whose values vary in size", "invalid_value", ) diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/codec/cast_value.py b/packages/zarr-metadata/src/zarr_metadata/v3/codec/cast_value.py index d5a018a1e6..921f05132f 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/codec/cast_value.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/codec/cast_value.py @@ -10,21 +10,16 @@ from typing_extensions import TypedDict from zarr_metadata._common import JSONValue -from zarr_metadata._json import ValidationProblem +from zarr_metadata._json import ValidationProblem, shown from zarr_metadata.v3._definition import ( + AcceptedField, Chunk, CodecDefinition, DataTypeField, Nested, fill_value_problems, - named_configuration, -) -from zarr_metadata.v3.codec._arithmetic import ( - COMPLEX, - FLOATING_POINT, - NOT_NUMBERS, - read_name, ) +from zarr_metadata.v3.codec._arithmetic import COMPLEX, FLOATING_POINT, NOT_NUMBERS CAST_VALUE_CODEC_NAME: Final = "cast_value" """The `name` field value of the `cast_value` codec.""" @@ -153,21 +148,20 @@ def _rules( models no real numbers is the one problem reported of it. """ target = nested.get(("data_type",)) - name = read_name(target) - if target is None or name is None: + if not isinstance(target, AcceptedField): return - written, _, _ = named_configuration(target.json) + name, written = target.definition.name, target.name if name in _NO_REAL_NUMBERS: yield ValidationProblem( ("data_type",), - f"expected a data type that models real numbers, got {written!r}", + f"expected a data type that models real numbers, got {shown(written)}", "invalid_value", ) return if configuration.get("out_of_range") == "wrap" and name in FLOATING_POINT: yield ValidationProblem( ("out_of_range",), - f"expected an integral data_type to wrap to, got {written!r}", + f"expected an integral data_type to wrap to, got {shown(written)}", "invalid_value", ) for at, scalar in _scalars(configuration, "target"): @@ -185,14 +179,12 @@ def _chunk_rules( so the data type it is handed is held to what the one it casts to is. """ source = chunk.data_type - name = read_name(source) - if source is None or name is None: + if not isinstance(source, AcceptedField): return - if name in _NO_REAL_NUMBERS: - written, _, _ = named_configuration(source.json) + if source.definition.name in _NO_REAL_NUMBERS: yield ValidationProblem( (), - f"expected a chunk of a data type that models real numbers, got {written!r}", + f"expected a chunk of a data type that models real numbers, got {shown(source.name)}", "invalid_value", ) return diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/codec/gzip.py b/packages/zarr-metadata/src/zarr_metadata/v3/codec/gzip.py index aa07d18421..66e535a165 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/codec/gzip.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/codec/gzip.py @@ -4,13 +4,12 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/codecs/gzip/index.html """ -from collections.abc import Iterator -from typing import Final, Literal, NotRequired +from typing import Annotated, Final, Literal, NotRequired +from annotated_types import Interval from typing_extensions import TypedDict -from zarr_metadata._json import ValidationProblem -from zarr_metadata.v3._definition import CodecDefinition, Nested +from zarr_metadata.v3._definition import CodecDefinition GZIP_CODEC_NAME: Final = "gzip" """The `name` field value of the `gzip` codec.""" @@ -33,7 +32,7 @@ class GzipCodecConfiguration(TypedDict, closed=True): https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/codecs/gzip/index.rst#L57-L66 """ - level: int + level: Annotated[int, Interval(ge=0, le=9)] class GzipCodecObject(TypedDict, closed=True): @@ -53,21 +52,11 @@ class GzipCodecObject(TypedDict, closed=True): """ -def _rules(configuration: GzipCodecConfiguration, nested: Nested) -> Iterator[ValidationProblem]: - """`level` is an integer from 0 to 9.""" - level = configuration["level"] - if not 0 <= level <= 9: - yield ValidationProblem( - ("level",), f"expected an integer in [0, 9], got {level}", "invalid_value" - ) - - GZIP_CODEC: Final = CodecDefinition( name=GZIP_CODEC_NAME, configuration=GzipCodecConfiguration, kind="bytes_bytes", size="dynamic", - rules=_rules, ) """The `gzip` codec.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/codec/scale_offset.py b/packages/zarr-metadata/src/zarr_metadata/v3/codec/scale_offset.py index e4807fe6fe..46af385734 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/codec/scale_offset.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/codec/scale_offset.py @@ -10,15 +10,15 @@ from typing_extensions import TypedDict from zarr_metadata._common import JSONValue -from zarr_metadata._json import ValidationProblem +from zarr_metadata._json import ValidationProblem, shown from zarr_metadata.v3._definition import ( + AcceptedField, Chunk, CodecDefinition, Nested, fill_value_problems, - named_configuration, ) -from zarr_metadata.v3.codec._arithmetic import NOT_NUMBERS, read_name +from zarr_metadata.v3.codec._arithmetic import NOT_NUMBERS SCALE_OFFSET_CODEC_NAME: Final = "scale_offset" """The `name` field value of the `scale_offset` codec.""" @@ -95,14 +95,12 @@ def _chunk_rules( A null is the rules' to refuse. """ source = chunk.data_type - name = read_name(source) - if source is None or name is None: + if not isinstance(source, AcceptedField): return - if name in NOT_NUMBERS: - written, _, _ = named_configuration(source.json) + if source.definition.name in NOT_NUMBERS: yield ValidationProblem( (), - f"expected a chunk of a data type with arithmetic, got {written!r}", + f"expected a chunk of a data type with arithmetic, got {shown(source.name)}", "invalid_value", ) return diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/codec/sharding_indexed.py b/packages/zarr-metadata/src/zarr_metadata/v3/codec/sharding_indexed.py index 98dd1560cf..e701c7d9de 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/codec/sharding_indexed.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/codec/sharding_indexed.py @@ -5,18 +5,19 @@ """ from collections.abc import Iterator, Mapping -from typing import Final, Literal, NotRequired +from typing import Annotated, Final, Literal, NotRequired, cast +from annotated_types import Ge from typing_extensions import TypedDict from zarr_metadata._json import ValidationProblem from zarr_metadata.v3._definition import ( + AcceptedField, Chunk, CodecDefinition, CodecField, Lengths, Nested, - Resolved, StaticCodecField, ) from zarr_metadata.v3.data_type.uint64 import UINT64_DATA_TYPE, UINT64_DATA_TYPE_NAME @@ -53,7 +54,7 @@ class ShardingIndexedCodecConfiguration(TypedDict, closed=True): https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/codecs/sharding-indexed/index.rst#L157-L161 """ - chunk_shape: tuple[int, ...] + chunk_shape: tuple[Annotated[int, Ge(1)], ...] codecs: tuple[CodecField, ...] index_codecs: tuple[StaticCodecField, ...] index_location: NotRequired[ShardingIndexLocation] @@ -73,22 +74,12 @@ class ShardingIndexedCodecObject(TypedDict, closed=True): The configuration has multiple required keys (`chunk_shape`, `codecs`, `index_codecs`), so only the object form is valid; the short-hand-name form is not permitted by the spec for this codec. - https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/codecs/sharding-indexed/index.rst#L141-L155 (required members) + https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/codecs/sharding-indexed/index.rst#L129-L138 (`chunk_shape`) + https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/codecs/sharding-indexed/index.rst#L141-L155 (`codecs` and `index_codecs`, the members the spec marks required) https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L1562-L1564 (short-hand names only "if no configuration metadata is required") """ -def _rules( - configuration: ShardingIndexedCodecConfiguration, nested: Nested -) -> Iterator[ValidationProblem]: - """Every inner chunk extent is at least 1.""" - for index, extent in enumerate(configuration["chunk_shape"]): - if extent < 1: - yield ValidationProblem( - ("chunk_shape", index), f"expected an integer >= 1, got {extent}", "invalid_value" - ) - - def _chunk_rules( configuration: ShardingIndexedCodecConfiguration, nested: Nested, chunk: Chunk ) -> Iterator[ValidationProblem]: @@ -125,7 +116,12 @@ def _chunk_rules( ) -_UINT64: Final = Resolved(UINT64_DATA_TYPE_NAME, "read", UINT64_DATA_TYPE, {}) +_UINT64: Final = AcceptedField( + json=UINT64_DATA_TYPE_NAME, + name=UINT64_DATA_TYPE_NAME, + definition=UINT64_DATA_TYPE, + configuration={}, +) """The data type of a shard index, which the spec fixes whatever the scope holds.""" @@ -170,12 +166,29 @@ def _per_shard(chunk_shape: tuple[int, ...], lengths: Lengths | None) -> Lengths ) +def _canonical( + configuration: ShardingIndexedCodecConfiguration, +) -> ShardingIndexedCodecConfiguration: + """Without an `index_location` of `end`, which is what an absent one means. + + "If the parameter is not present, the value defaults to `end`" + (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/codecs/sharding-indexed/index.rst#L157-L161), + so the two spellings are one codec. + """ + if configuration.get("index_location") != "end": + return configuration + return cast( + "ShardingIndexedCodecConfiguration", + {key: value for key, value in configuration.items() if key != "index_location"}, + ) + + SHARDING_INDEXED_CODEC: Final = CodecDefinition( name=SHARDING_INDEXED_CODEC_NAME, configuration=ShardingIndexedCodecConfiguration, + canonical=_canonical, kind="array_bytes", size="dynamic", - rules=_rules, chunk_rules=_chunk_rules, pipelines=_pipelines, ) diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/codec/transpose.py b/packages/zarr-metadata/src/zarr_metadata/v3/codec/transpose.py index 2fb88338f9..83805c4c1a 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/codec/transpose.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/codec/transpose.py @@ -9,7 +9,7 @@ from typing_extensions import TypedDict -from zarr_metadata._json import ValidationProblem +from zarr_metadata._json import ValidationProblem, shown from zarr_metadata.v3._definition import Chunk, CodecDefinition, Nested TRANSPOSE_CODEC_NAME: Final = "transpose" @@ -60,7 +60,7 @@ def _rules( if sorted(order) != list(range(len(order))): yield ValidationProblem( ("order",), - f"expected a permutation of 0..{len(order) - 1}, got {order!r}", + f"expected a permutation of 0..{len(order) - 1}, got {shown(order)}", "invalid_value", ) diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/codec/zstd.py b/packages/zarr-metadata/src/zarr_metadata/v3/codec/zstd.py index efeefb7edb..55cd471dae 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/codec/zstd.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/codec/zstd.py @@ -6,13 +6,12 @@ proposed the codec, was never merged). """ -from collections.abc import Iterator -from typing import Final, Literal, NotRequired +from typing import Annotated, Final, Literal, NotRequired, cast +from annotated_types import Interval from typing_extensions import TypedDict -from zarr_metadata._json import ValidationProblem -from zarr_metadata.v3._definition import CodecDefinition, Nested +from zarr_metadata.v3._definition import CodecDefinition ZSTD_CODEC_NAME: Final = "zstd" """The `name` field value of the `zstd` codec.""" @@ -20,6 +19,12 @@ ZstdCodecName = Literal["zstd"] """Literal type of the `name` field of the `zstd` codec.""" +ZSTD_MIN_LEVEL: Final = -131072 +"""The lowest `level` zstd accepts: ZSTD_minCLevel(), -(1 << 17).""" + +ZSTD_MAX_LEVEL: Final = 22 +"""The highest `level` zstd accepts: ZSTD_maxCLevel().""" + class ZstdCodecConfiguration(TypedDict, closed=True): """ @@ -30,7 +35,7 @@ class ZstdCodecConfiguration(TypedDict, closed=True): https://github.com/zarr-developers/zarr-extensions/blob/4da7b37a84f76e660902f6d3de3eaef0e0febae6/codecs/zstd/README.md#L9-L19 """ - level: int + level: Annotated[int, Interval(ge=ZSTD_MIN_LEVEL, le=ZSTD_MAX_LEVEL)] checksum: NotRequired[bool] @@ -51,30 +56,28 @@ class ZstdCodecObject(TypedDict, closed=True): https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L1562-L1564 (short-hand names only "if no configuration metadata is required") """ -ZSTD_MIN_LEVEL: Final = -131072 -"""The lowest `level` zstd accepts: ZSTD_minCLevel(), -(1 << 17).""" - -ZSTD_MAX_LEVEL: Final = 22 -"""The highest `level` zstd accepts: ZSTD_maxCLevel().""" +def _canonical(configuration: ZstdCodecConfiguration) -> ZstdCodecConfiguration: + """Without a `checksum` of `false`, which is what an absent one means. -def _rules(configuration: ZstdCodecConfiguration, nested: Nested) -> Iterator[ValidationProblem]: - """`level` is one zstd accepts.""" - level = configuration["level"] - if not ZSTD_MIN_LEVEL <= level <= ZSTD_MAX_LEVEL: - yield ValidationProblem( - ("level",), - f"expected an integer in [{ZSTD_MIN_LEVEL}, {ZSTD_MAX_LEVEL}], got {level}", - "invalid_value", - ) + The spec says of `checksum` that it "should be omitted if false" + (https://github.com/zarr-developers/zarr-extensions/blob/4da7b37a84f76e660902f6d3de3eaef0e0febae6/codecs/zstd/README.md?plain=1#L17-L19), + so the two spellings are one codec. + """ + if configuration.get("checksum") is not False: + return configuration + return cast( + "ZstdCodecConfiguration", + {key: value for key, value in configuration.items() if key != "checksum"}, + ) ZSTD_CODEC: Final = CodecDefinition( name=ZSTD_CODEC_NAME, configuration=ZstdCodecConfiguration, + canonical=_canonical, kind="bytes_bytes", size="dynamic", - rules=_rules, ) """The `zstd` codec.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_byte.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_byte.py new file mode 100644 index 0000000000..43a2aaed66 --- /dev/null +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_byte.py @@ -0,0 +1,15 @@ +"""A byte value, as the fill values of raw bits and of `bytes` hold them. + +Each is an integer in `[0, 255]` +(https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/data-types/index.rst#L97-L99; +the `bytes` type's is https://github.com/zarr-developers/zarr-extensions/blob/4da7b37a84f76e660902f6d3de3eaef0e0febae6/data-types/bytes/README.md?plain=1#L8). +""" + +from typing import Annotated + +from annotated_types import Interval + +ByteValue = Annotated[int, Interval(ge=0, le=255)] +"""One byte of a fill value: a JSON integer in `[0, 255]`.""" + +__all__ = ["ByteValue"] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_float.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_float.py index 3b81bd7a7c..595ee23bf4 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_float.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_float.py @@ -9,11 +9,15 @@ (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/data-types/index.rst#L88-L91). """ -import functools +import dataclasses +import struct from collections.abc import Callable, Iterable, Iterator -from typing import Any, Literal, get_args +from dataclasses import dataclass +from decimal import ROUND_HALF_EVEN, Decimal, localcontext +from typing import Any, Final, Literal, get_args -from zarr_metadata._json import ValidationProblem +from zarr_metadata._common import JSONValue +from zarr_metadata._json import ValidationProblem, choices, shown from zarr_metadata.v3._definition import EmptyConfiguration, Nested FloatSpecialFillValue = Literal["NaN", "Infinity", "-Infinity"] @@ -28,30 +32,149 @@ def float_fill_value_rules( A number takes any value, since a reader rounds it. A string is one of the named values, or a hex string of the type's own width: the type-checked shape takes any string there, and `hex_form` raises - `ValueError` for one that is not. A partial application of a - module-level function, so a definition holding it pickles. + `ValueError` for one that is not. """ - return functools.partial(_float_fill_value, name, hex_form) - - -def _float_fill_value( - name: str, - hex_form: Callable[[str], str], - configuration: EmptyConfiguration, - nested: Nested, - value: float | str, -) -> Iterator[ValidationProblem]: - if not isinstance(value, str) or value in get_args(FloatSpecialFillValue): - return + return _FloatFillValue(name, hex_form) + + +@dataclass(frozen=True, slots=True) +class _FloatFillValue: + """A float fill value: a value rather than a closure, so a definition holding it is equal to itself after a pickle or a deep copy.""" + + name: str + hex_form: Callable[[str], str] + + def __call__( + self, configuration: EmptyConfiguration, nested: Nested, value: float | str + ) -> Iterator[ValidationProblem]: + if not isinstance(value, str) or value in get_args(FloatSpecialFillValue): + return + try: + self.hex_form(value) + except ValueError: + yield ValidationProblem( + (), + f"expected a number, {choices(get_args(FloatSpecialFillValue))}, or a " + f"{self.name} hex string, got {shown(value)}", + "invalid_value", + ) + + +FloatWidth = Literal[16, 32, 64] +"""The widths, in bits, of the IEEE 754 binary formats the float types store.""" + +_FORMATS: Final[dict[FloatWidth, tuple[str, int]]] = {16: ("e", 10), 32: ("f", 23), 64: ("d", 52)} +"""Each width's `struct` format, and how many bits of it hold the fraction.""" + + +def float_fill_value_canonical( + width: FloatWidth, +) -> Callable[[EmptyConfiguration, Nested, float | str], float | str]: + """The canonical spelling of a fill value of the floating-point type `width` bits wide. + + A fill value spells a value of the type, as `float_bits` reads one: a + number, a named value, or the value's bits. Its canonical spelling is + the named value for an infinity, or for the NaN the spec names + `"NaN"`; the hex string of its bits, in lower case, for any other NaN; + and, for any other value, the shortest number that rounds to it, the + nearest of those, as numpy and `repr` spell one -- `0.1` for the + `float32` nearest `0.1` -- `-0.0`, a value of its own, among them. + """ + return _FloatCanonical(width) + + +@dataclass(frozen=True, slots=True) +class _FloatCanonical: + """A float fill value's canonical spelling: a value rather than a closure, as `_FloatFillValue` is.""" + + width: FloatWidth + + def __call__( + self, configuration: EmptyConfiguration, nested: Nested, value: float | str + ) -> float | str: + return _spelled(float_bits(value, self.width), self.width) + + +def float_bits(value: float | str, width: FloatWidth) -> int: + """The bits of the value `value`, a float fill value the rules allow, spells in the type `width` bits wide. + + A number is read as a float64, as a JSON parser reads one -- an integer + of more digits than a float64 holds is rounded to one -- and rounded to + the nearest value the type represents, ties to even, and to an infinity + past the largest: as zarrs reads one, `as_f64` then `as f32` + (https://github.com/zarrs/zarrs/blob/8d68f8522b382d050b768f84bce64c2935de4523/zarrs_metadata/src/v3/array/fill_value.rs#L160), + as tensorstore does, `static_cast` of `get()` + (https://github.com/google/tensorstore/blob/692d2798c51a76d2eed0b4aee85cad5fd4be950a/tensorstore/driver/zarr3/metadata.cc#L165), + and as numpy casts a float64. So an integer and a number with a + fraction that read as one float64 spell one value; and an integer past + 2**53 whose float64 sits halfway between two values of the type -- + 2**60 + 2**36 + 1, for float32 -- is the even one, 2**60, as those + readers store it, not the value nearest the integer itself. + """ + code, fraction = _FORMATS[width] + exponent = width - 1 - fraction + infinity = ((1 << exponent) - 1) << fraction + sign = 1 << (width - 1) + if isinstance(value, str): + named = { + "NaN": infinity | (1 << (fraction - 1)), + "Infinity": infinity, + "-Infinity": sign | infinity, + } + return named[value] if value in named else int(value, 16) + try: + held = float(value) + except OverflowError: + # An integer past the largest float64 reads as an infinity. + return (sign if value < 0 else 0) | infinity try: - hex_form(value) - except ValueError: - yield ValidationProblem( - (), - f"expected a number, one of {get_args(FloatSpecialFillValue)!r}, or a {name} " - f"hex string, got {value!r}", - "invalid_value", - ) + return int.from_bytes(struct.pack(f">{code}", held), "big") + except OverflowError: + # `struct` refuses a float64 that rounds past the type's largest value. + return (sign if held < 0 else 0) | infinity + + +def _spelled(bits: int, width: FloatWidth) -> float | str: + """The canonical spelling of the value whose bits, in the type `width` bits wide, are `bits`.""" + code, fraction = _FORMATS[width] + exponent = width - 1 - fraction + infinity = ((1 << exponent) - 1) << fraction + sign = 1 << (width - 1) + if bits & infinity == infinity: + if bits & ((1 << fraction) - 1) == 0: + return "-Infinity" if bits & sign else "Infinity" + if bits == infinity | (1 << (fraction - 1)): + return "NaN" + return f"0x{bits:0{width // 4}x}" + value: float = struct.unpack(f">{code}", bits.to_bytes(width // 8, "big"))[0] + if width == 64 or value == 0: + # A float64 is its own shortest spelling, as `repr` writes it, and + # so is a zero of either sign. + return value + return _shortest(value, bits, width) + + +def _shortest(value: float, bits: int, width: FloatWidth) -> float: + """The shortest number that rounds to `value`, whose bits in the type `width` bits wide are `bits`: of the fewest significant digits, the nearest to it. + + Of each number of digits the nearest is tried, and then the one past + `value` from it: at a power of two the numbers that round to it reach + twice as far above it as below, so the nearest may miss where the next + one does not. + """ + exact = Decimal(value) + for digits in range(1, 18): + with localcontext() as context: + context.prec = digits + context.rounding = ROUND_HALF_EVEN + nearest = +exact + step = Decimal((0, (1,), nearest.adjusted() - digits + 1)) + beyond = nearest + step if nearest < exact else nearest - step + for candidate in (nearest, beyond): + if float_bits(float(candidate), width) == bits: + return float(candidate) + # Seventeen significant digits tell every float64 from every other. + return value def complex_fill_value_rules( @@ -63,18 +186,54 @@ def complex_fill_value_rules( `component` is the fill value rules of the component's own float type. """ - return functools.partial(_complex_fill_value, component) + return _ComplexFillValue(component) -def _complex_fill_value( - component: Callable[[EmptyConfiguration, Nested, Any], Iterable[ValidationProblem]], - configuration: EmptyConfiguration, - nested: Nested, - value: tuple[float | str, float | str], -) -> Iterator[ValidationProblem]: - for index, part in enumerate(value): - for found in component(configuration, nested, part): - yield ValidationProblem((index, *found.loc), found.message, found.kind) +@dataclass(frozen=True, slots=True) +class _ComplexFillValue: + """A complex fill value, each component judged at its index: a value rather than a closure, as `_FloatFillValue` is.""" + + component: Callable[[EmptyConfiguration, Nested, Any], Iterable[ValidationProblem]] + + def __call__( + self, + configuration: EmptyConfiguration, + nested: Nested, + value: tuple[float | str, float | str], + ) -> Iterator[ValidationProblem]: + for index, part in enumerate(value): + for found in self.component(configuration, nested, part): + yield dataclasses.replace(found, loc=(index, *found.loc)) + + +def complex_fill_value_canonical( + component: Callable[[EmptyConfiguration, Nested, Any], JSONValue], +) -> Callable[[EmptyConfiguration, Nested, tuple[float | str, float | str]], JSONValue]: + """The canonical spelling of a complex fill value: each component in the canonical spelling `component`, its float type's, gives it.""" + return _ComplexCanonical(component) + + +@dataclass(frozen=True, slots=True) +class _ComplexCanonical: + """A complex fill value's canonical spelling: a value rather than a closure, as `_FloatFillValue` is.""" + + component: Callable[[EmptyConfiguration, Nested, Any], JSONValue] + + def __call__( + self, + configuration: EmptyConfiguration, + nested: Nested, + value: tuple[float | str, float | str], + ) -> JSONValue: + return tuple(self.component(configuration, nested, part) for part in value) -__all__ = ["FloatSpecialFillValue", "complex_fill_value_rules", "float_fill_value_rules"] +__all__ = [ + "FloatSpecialFillValue", + "FloatWidth", + "complex_fill_value_canonical", + "complex_fill_value_rules", + "float_bits", + "float_fill_value_canonical", + "float_fill_value_rules", +] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_integer.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_integer.py deleted file mode 100644 index 4bf1005fd5..0000000000 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_integer.py +++ /dev/null @@ -1,44 +0,0 @@ -"""The fill value rules integers share: a JSON integer within a range. - -The integer data types take one within their own range, and the byte -values of `r*` and `bytes` fill values are integers in `[0, 255]` -(https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/data-types/index.rst#L59-L61). -""" - -import functools -from collections.abc import Callable, Iterator - -from zarr_metadata._json import ValidationProblem -from zarr_metadata.v3._definition import EmptyConfiguration, Nested - - -def integer_fill_value_rules( - low: int, high: int -) -> Callable[[EmptyConfiguration, Nested, int], Iterator[ValidationProblem]]: - """The fill value rules of an integer data type whose range is `[low, high]`. - - A partial application of a module-level function, so a definition - holding it pickles. - """ - return functools.partial(_in_range, low, high) - - -def _in_range( - low: int, high: int, configuration: EmptyConfiguration, nested: Nested, value: int -) -> Iterator[ValidationProblem]: - if not low <= value <= high: - yield ValidationProblem( - (), f"expected an integer in [{low}, {high}], got {value}", "invalid_value" - ) - - -def byte_value_problems(values: tuple[int, ...]) -> Iterator[ValidationProblem]: - """Each of `values` that is not a byte, an integer in `[0, 255]`, located at its index.""" - for index, value in enumerate(values): - if not 0 <= value <= 255: - yield ValidationProblem( - (index,), f"expected an integer in [0, 255], got {value}", "invalid_value" - ) - - -__all__ = ["byte_value_problems", "integer_fill_value_rules"] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_numpy_time.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_numpy_time.py index 7023495b93..de5c0c7d09 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_numpy_time.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/_numpy_time.py @@ -1,21 +1,47 @@ -"""What the two numpy time types share: a unit, and how many of it one tick is. +"""What the two numpy time types share: how many of a unit one tick is, a count of ticks, and how their values are stored. -Both types' configurations have these two members and the one rule on -them, so the rule is written here once and neither sibling imports it -from the other. So is how their values are stored. +Both types' configurations hold a `scale_factor` of one type, and both +types' fill values count ticks of one width, so each is written here once +and neither sibling imports it from the other. So is how their values are +stored. """ -from collections.abc import Iterator -from typing import Final +from typing import Annotated, Final -from typing_extensions import ReadOnly, TypedDict +from annotated_types import Interval -from zarr_metadata._json import ValidationProblem -from zarr_metadata.v3._definition import Nested, multi_byte +from zarr_metadata.v3._definition import multi_byte NUMPY_TIME_MAX_SCALE_FACTOR: Final = 2**31 - 1 """The largest `scale_factor` numpy stores: the field is a signed int32.""" +NumpyTimeScaleFactor = Annotated[int, Interval(ge=1, le=NUMPY_TIME_MAX_SCALE_FACTOR)] +"""How many of the unit one tick is: a positive int32.""" + +NumpyTimeTicks = Annotated[int, Interval(ge=-(2**63), le=2**63 - 1)] +"""A count of ticks, as a fill value writes one: a signed 64-bit integer. + +https://github.com/zarr-developers/zarr-extensions/blob/4da7b37a84f76e660902f6d3de3eaef0e0febae6/data-types/numpy.datetime64/README.md?plain=1#L109-L112 +""" + +NOT_A_TIME_TICKS: Final = -(2**63) +"""The tick count `NaT` is stored as, which a fill value may write for it. + +"`"fill_value": "NaT"` and `"fill_value": -9223372036854775808` should +be treated as equivalent representations of the same scalar value" +(https://github.com/zarr-developers/zarr-extensions/blob/4da7b37a84f76e660902f6d3de3eaef0e0febae6/data-types/numpy.datetime64/README.md?plain=1#L114-L116); +`numpy.timedelta64` says the same +(https://github.com/zarr-developers/zarr-extensions/blob/4da7b37a84f76e660902f6d3de3eaef0e0febae6/data-types/numpy.timedelta64/README.md?plain=1#L117-L119). +""" + + +def numpy_time_fill_value_canonical( + configuration: object, nested: object, value: int | str +) -> int | str: + """The canonical spelling of a numpy time fill value: `"NaT"` for `NaT`, however it is written, and any other count of ticks as written.""" + return "NaT" if value in ("NaT", NOT_A_TIME_TICKS) else value + + numpy_time_storage: Final = multi_byte """Signed 64-bit integers, in the byte order the codecs say. @@ -29,43 +55,11 @@ """ -class NumpyTimeConfiguration(TypedDict): - """The members both numpy time types' configurations have, read-only so either type fits.""" - - unit: ReadOnly[str] - scale_factor: ReadOnly[int] - - -def numpy_time_rules( - configuration: NumpyTimeConfiguration, nested: Nested -) -> Iterator[ValidationProblem]: - """`scale_factor` is a positive int32.""" - scale_factor = configuration["scale_factor"] - if not 1 <= scale_factor <= NUMPY_TIME_MAX_SCALE_FACTOR: - yield ValidationProblem( - ("scale_factor",), - f"expected an integer in [1, {NUMPY_TIME_MAX_SCALE_FACTOR}], got {scale_factor}", - "invalid_value", - ) - - -def numpy_time_fill_value_rules( - configuration: NumpyTimeConfiguration, nested: Nested, value: int | str -) -> Iterator[ValidationProblem]: - """An integer fill value is a signed 64-bit one; `"NaT"` is the other form, which the shape admits. - - https://github.com/zarr-developers/zarr-extensions/blob/4da7b37a84f76e660902f6d3de3eaef0e0febae6/data-types/numpy.datetime64/README.md?plain=1#L109-L112 - """ - if isinstance(value, int) and not -(2**63) <= value <= 2**63 - 1: - yield ValidationProblem( - (), f"expected a signed 64-bit integer or 'NaT', got {value}", "invalid_value" - ) - - __all__ = [ + "NOT_A_TIME_TICKS", "NUMPY_TIME_MAX_SCALE_FACTOR", - "NumpyTimeConfiguration", - "numpy_time_fill_value_rules", - "numpy_time_rules", + "NumpyTimeScaleFactor", + "NumpyTimeTicks", + "numpy_time_fill_value_canonical", "numpy_time_storage", ] diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/bytes.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/bytes.py index 530b88a5c2..aa83447ef3 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/bytes.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/bytes.py @@ -4,18 +4,19 @@ See https://github.com/zarr-developers/zarr-extensions/blob/4da7b37a84f76e660902f6d3de3eaef0e0febae6/data-types/bytes/README.md """ +import base64 import re from collections.abc import Iterator from typing import Final, Literal, NewType -from zarr_metadata._json import ValidationProblem +from zarr_metadata._json import ValidationProblem, shown from zarr_metadata.v3._definition import ( DataTypeDefinition, EmptyConfiguration, Nested, variable_length, ) -from zarr_metadata.v3.data_type._integer import byte_value_problems +from zarr_metadata.v3.data_type._byte import ByteValue BYTES_DATA_TYPE_NAME: Final = "bytes" """The `data_type` value for the variable-length `bytes` type.""" @@ -41,7 +42,7 @@ def base64_bytes(value: str) -> Base64Bytes: return Base64Bytes(value) -BytesFillValue = tuple[int, ...] | Base64Bytes +BytesFillValue = tuple[ByteValue, ...] | Base64Bytes """Permitted JSON shape of the `fill_value` field for `bytes`. Either a JSON array of integers in `[0, 255]` (one per byte), or a @@ -52,23 +53,37 @@ def base64_bytes(value: str) -> Base64Bytes: def _fill_value_rules( configuration: EmptyConfiguration, nested: Nested, value: BytesFillValue ) -> Iterator[ValidationProblem]: - """Integers in `[0, 255]`, or a string of standard-alphabet base64.""" + """A string of standard-alphabet base64, when it is not byte values.""" if not isinstance(value, str): - yield from byte_value_problems(value) return try: base64_bytes(value) except ValueError: yield ValidationProblem( - (), f"expected standard-alphabet base64, got {value!r}", "invalid_value" + (), f"expected standard-alphabet base64, got {shown(value)}", "invalid_value" ) +def _fill_value_canonical( + configuration: EmptyConfiguration, nested: Nested, value: BytesFillValue +) -> str: + """The bytes, however written, as the base64 that encodes them. + + The string "is more compact" + (https://github.com/zarr-developers/zarr-extensions/blob/4da7b37a84f76e660902f6d3de3eaef0e0febae6/data-types/bytes/README.md?plain=1#L11), + and encoding the bytes again spells them one way: `"QR=="` decodes to + the one byte `"QQ=="` does. + """ + data = base64.b64decode(value) if isinstance(value, str) else bytes(value) + return base64.b64encode(data).decode("ascii") + + BYTES_DATA_TYPE: Final = DataTypeDefinition( name=BYTES_DATA_TYPE_NAME, configuration=EmptyConfiguration, fill_value=BytesFillValue, fill_value_rules=_fill_value_rules, + fill_value_canonical=_fill_value_canonical, storage=variable_length, ) """The `bytes` data type: a bare name, with nothing to configure. diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/complex128.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/complex128.py index 509b12f30a..bf3d55de87 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/complex128.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/complex128.py @@ -7,7 +7,10 @@ from typing import Final, Literal from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._float import complex_fill_value_rules +from zarr_metadata.v3.data_type._float import ( + complex_fill_value_canonical, + complex_fill_value_rules, +) from zarr_metadata.v3.data_type.float64 import FLOAT64_DATA_TYPE, Float64FillValue COMPLEX128_DATA_TYPE_NAME: Final = "complex128" @@ -36,6 +39,7 @@ configuration=EmptyConfiguration, fill_value=Complex128FillValue, fill_value_rules=complex_fill_value_rules(FLOAT64_DATA_TYPE.fill_value_rules), + fill_value_canonical=complex_fill_value_canonical(FLOAT64_DATA_TYPE.fill_value_canonical), storage=multi_byte, ) """The `complex128` data type: a bare name, with nothing to configure; its fill value a pair of `float64` components.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/complex64.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/complex64.py index ceb1576473..5ddd0559f8 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/complex64.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/complex64.py @@ -7,7 +7,10 @@ from typing import Final, Literal from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._float import complex_fill_value_rules +from zarr_metadata.v3.data_type._float import ( + complex_fill_value_canonical, + complex_fill_value_rules, +) from zarr_metadata.v3.data_type.float32 import FLOAT32_DATA_TYPE, Float32FillValue COMPLEX64_DATA_TYPE_NAME: Final = "complex64" @@ -36,6 +39,7 @@ configuration=EmptyConfiguration, fill_value=Complex64FillValue, fill_value_rules=complex_fill_value_rules(FLOAT32_DATA_TYPE.fill_value_rules), + fill_value_canonical=complex_fill_value_canonical(FLOAT32_DATA_TYPE.fill_value_canonical), storage=multi_byte, ) """The `complex64` data type: a bare name, with nothing to configure; its fill value a pair of `float32` components.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float16.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float16.py index 33ba182dc4..1a53e8bca1 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float16.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float16.py @@ -8,7 +8,11 @@ from typing import Final, Literal, NewType from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._float import FloatSpecialFillValue, float_fill_value_rules +from zarr_metadata.v3.data_type._float import ( + FloatSpecialFillValue, + float_fill_value_canonical, + float_fill_value_rules, +) FLOAT16_DATA_TYPE_NAME: Final = "float16" """The `data_type` value for the `float16` type.""" @@ -69,6 +73,7 @@ def hex_float16(value: str) -> HexFloat16: configuration=EmptyConfiguration, fill_value=Float16FillValue, fill_value_rules=float_fill_value_rules("float16", hex_float16), + fill_value_canonical=float_fill_value_canonical(16), storage=multi_byte, ) """The `float16` data type: a bare name, with nothing to configure; its fill value a number, a named value or a hex string.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float32.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float32.py index 7eb923d581..35e9706ae0 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float32.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float32.py @@ -8,7 +8,11 @@ from typing import Final, Literal, NewType from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._float import FloatSpecialFillValue, float_fill_value_rules +from zarr_metadata.v3.data_type._float import ( + FloatSpecialFillValue, + float_fill_value_canonical, + float_fill_value_rules, +) FLOAT32_DATA_TYPE_NAME: Final = "float32" """The `data_type` value for the `float32` type.""" @@ -69,6 +73,7 @@ def hex_float32(value: str) -> HexFloat32: configuration=EmptyConfiguration, fill_value=Float32FillValue, fill_value_rules=float_fill_value_rules("float32", hex_float32), + fill_value_canonical=float_fill_value_canonical(32), storage=multi_byte, ) """The `float32` data type: a bare name, with nothing to configure; its fill value a number, a named value or a hex string.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float64.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float64.py index 0363439cfa..a086ed50a5 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float64.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/float64.py @@ -8,7 +8,11 @@ from typing import Final, Literal, NewType from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._float import FloatSpecialFillValue, float_fill_value_rules +from zarr_metadata.v3.data_type._float import ( + FloatSpecialFillValue, + float_fill_value_canonical, + float_fill_value_rules, +) FLOAT64_DATA_TYPE_NAME: Final = "float64" """The `data_type` value for the `float64` type.""" @@ -70,6 +74,7 @@ def hex_float64(value: str) -> HexFloat64: configuration=EmptyConfiguration, fill_value=Float64FillValue, fill_value_rules=float_fill_value_rules("float64", hex_float64), + fill_value_canonical=float_fill_value_canonical(64), storage=multi_byte, ) """The `float64` data type: a bare name, with nothing to configure; its fill value a number, a named value or a hex string.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int16.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int16.py index 22ff1fd8a4..adf03467b2 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int16.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int16.py @@ -4,10 +4,11 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/data-types/index.html """ -from typing import Final, Literal +from typing import Annotated, Final, Literal + +from annotated_types import Interval from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._integer import integer_fill_value_rules INT16_DATA_TYPE_NAME: Final = "int16" """The `data_type` value for the `int16` type.""" @@ -15,7 +16,7 @@ Int16DataTypeName = Literal["int16"] """Literal type of the `data_type` field for `int16`.""" -Int16FillValue = int +Int16FillValue = Annotated[int, Interval(ge=-(2**15), le=2**15 - 1)] """Permitted JSON shape of the `fill_value` field for `int16`: a JSON integer in [-32768, 32767].""" @@ -23,7 +24,6 @@ name=INT16_DATA_TYPE_NAME, configuration=EmptyConfiguration, fill_value=Int16FillValue, - fill_value_rules=integer_fill_value_rules(-(2**15), 2**15 - 1), storage=multi_byte, ) """The `int16` data type: a bare name, with nothing to configure; its fill value an integer in its range.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int32.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int32.py index 98a039a5d9..6a87f44476 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int32.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int32.py @@ -4,10 +4,11 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/data-types/index.html """ -from typing import Final, Literal +from typing import Annotated, Final, Literal + +from annotated_types import Interval from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._integer import integer_fill_value_rules INT32_DATA_TYPE_NAME: Final = "int32" """The `data_type` value for the `int32` type.""" @@ -15,7 +16,7 @@ Int32DataTypeName = Literal["int32"] """Literal type of the `data_type` field for `int32`.""" -Int32FillValue = int +Int32FillValue = Annotated[int, Interval(ge=-(2**31), le=2**31 - 1)] """Permitted JSON shape of the `fill_value` field for `int32`: a JSON integer in [-2**31, 2**31 - 1].""" @@ -23,7 +24,6 @@ name=INT32_DATA_TYPE_NAME, configuration=EmptyConfiguration, fill_value=Int32FillValue, - fill_value_rules=integer_fill_value_rules(-(2**31), 2**31 - 1), storage=multi_byte, ) """The `int32` data type: a bare name, with nothing to configure; its fill value an integer in its range.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int64.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int64.py index b49d5021b7..707bf88a90 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int64.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int64.py @@ -4,10 +4,11 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/data-types/index.html """ -from typing import Final, Literal +from typing import Annotated, Final, Literal + +from annotated_types import Interval from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._integer import integer_fill_value_rules INT64_DATA_TYPE_NAME: Final = "int64" """The `data_type` value for the `int64` type.""" @@ -15,7 +16,7 @@ Int64DataTypeName = Literal["int64"] """Literal type of the `data_type` field for `int64`.""" -Int64FillValue = int +Int64FillValue = Annotated[int, Interval(ge=-(2**63), le=2**63 - 1)] """Permitted JSON shape of the `fill_value` field for `int64`: a JSON integer in [-2**63, 2**63 - 1].""" @@ -23,7 +24,6 @@ name=INT64_DATA_TYPE_NAME, configuration=EmptyConfiguration, fill_value=Int64FillValue, - fill_value_rules=integer_fill_value_rules(-(2**63), 2**63 - 1), storage=multi_byte, ) """The `int64` data type: a bare name, with nothing to configure; its fill value an integer in its range.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int8.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int8.py index 7748a491f8..443b5f8a90 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int8.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/int8.py @@ -4,10 +4,11 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/data-types/index.html """ -from typing import Final, Literal +from typing import Annotated, Final, Literal + +from annotated_types import Interval from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, single_byte -from zarr_metadata.v3.data_type._integer import integer_fill_value_rules INT8_DATA_TYPE_NAME: Final = "int8" """The `data_type` value for the `int8` type.""" @@ -15,7 +16,7 @@ Int8DataTypeName = Literal["int8"] """Literal type of the `data_type` field for `int8`.""" -Int8FillValue = int +Int8FillValue = Annotated[int, Interval(ge=-(2**7), le=2**7 - 1)] """Permitted JSON shape of the `fill_value` field for `int8`: a JSON integer in [-128, 127].""" @@ -23,7 +24,6 @@ name=INT8_DATA_TYPE_NAME, configuration=EmptyConfiguration, fill_value=Int8FillValue, - fill_value_rules=integer_fill_value_rules(-(2**7), 2**7 - 1), storage=single_byte, ) """The `int8` data type: a bare name, with nothing to configure; its fill value an integer in its range.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/numpy_datetime64.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/numpy_datetime64.py index f8b2e21280..4464c93d4d 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/numpy_datetime64.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/numpy_datetime64.py @@ -10,8 +10,9 @@ from zarr_metadata.v3._definition import DataTypeDefinition from zarr_metadata.v3.data_type._numpy_time import ( - numpy_time_fill_value_rules, - numpy_time_rules, + NumpyTimeScaleFactor, + NumpyTimeTicks, + numpy_time_fill_value_canonical, numpy_time_storage, ) @@ -40,7 +41,7 @@ class NumpyDatetime64Configuration(TypedDict, closed=True): """ unit: ReadOnly[NumpyTimeUnit] - scale_factor: ReadOnly[int] + scale_factor: ReadOnly[NumpyTimeScaleFactor] class NumpyDatetime64(TypedDict, closed=True): @@ -51,7 +52,7 @@ class NumpyDatetime64(TypedDict, closed=True): must_understand: NotRequired[bool] -NumpyDatetime64FillValue = int | Literal["NaT"] +NumpyDatetime64FillValue = NumpyTimeTicks | Literal["NaT"] """Permitted JSON shape of the `fill_value` field for `numpy.datetime64`. Either a JSON integer (count of `unit * scale_factor` since the epoch), @@ -61,9 +62,8 @@ class NumpyDatetime64(TypedDict, closed=True): NUMPY_DATETIME64_DATA_TYPE: Final = DataTypeDefinition( name=NUMPY_DATETIME64_DATA_TYPE_NAME, configuration=NumpyDatetime64Configuration, - rules=numpy_time_rules, fill_value=NumpyDatetime64FillValue, - fill_value_rules=numpy_time_fill_value_rules, + fill_value_canonical=numpy_time_fill_value_canonical, storage=numpy_time_storage, ) """The `numpy.datetime64` data type.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/numpy_timedelta64.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/numpy_timedelta64.py index 5b0ee23626..0aa75d3559 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/numpy_timedelta64.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/numpy_timedelta64.py @@ -10,8 +10,9 @@ from zarr_metadata.v3._definition import DataTypeDefinition from zarr_metadata.v3.data_type._numpy_time import ( - numpy_time_fill_value_rules, - numpy_time_rules, + NumpyTimeScaleFactor, + NumpyTimeTicks, + numpy_time_fill_value_canonical, numpy_time_storage, ) @@ -59,7 +60,7 @@ class NumpyTimedelta64Configuration(TypedDict, closed=True): """ unit: ReadOnly[NumpyTimeUnit] - scale_factor: ReadOnly[int] + scale_factor: ReadOnly[NumpyTimeScaleFactor] class NumpyTimedelta64(TypedDict, closed=True): @@ -70,7 +71,7 @@ class NumpyTimedelta64(TypedDict, closed=True): must_understand: NotRequired[bool] -NumpyTimedelta64FillValue = int | Literal["NaT"] +NumpyTimedelta64FillValue = NumpyTimeTicks | Literal["NaT"] """Permitted JSON shape of the `fill_value` field for `numpy.timedelta64`. Either a JSON integer (a count of `unit * scale_factor`), or the string @@ -80,9 +81,8 @@ class NumpyTimedelta64(TypedDict, closed=True): NUMPY_TIMEDELTA64_DATA_TYPE: Final = DataTypeDefinition( name=NUMPY_TIMEDELTA64_DATA_TYPE_NAME, configuration=NumpyTimedelta64Configuration, - rules=numpy_time_rules, fill_value=NumpyTimedelta64FillValue, - fill_value_rules=numpy_time_fill_value_rules, + fill_value_canonical=numpy_time_fill_value_canonical, storage=numpy_time_storage, ) """The `numpy.timedelta64` data type.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/raw.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/raw.py index e23e523f8c..b7475cb2f2 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/raw.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/raw.py @@ -21,7 +21,7 @@ Nested, StorageClass, ) -from zarr_metadata.v3.data_type._integer import byte_value_problems +from zarr_metadata.v3.data_type._byte import ByteValue RawBytesDataTypeName = NewType("RawBytesDataTypeName", str) """A spec-conformant `r` raw-bytes name (e.g. `"r8"`, `"r16"`). @@ -46,7 +46,7 @@ def raw_bytes_dtype_name(value: str) -> RawBytesDataTypeName: return RawBytesDataTypeName(value) -RawBytesFillValue = tuple[int, ...] +RawBytesFillValue = tuple[ByteValue, ...] """Permitted JSON shape of the `fill_value` field for `r`. A JSON array of N/8 integers in `[0, 255]` (one per byte). @@ -78,7 +78,7 @@ def _rules(configuration: RawBytesConfiguration, nested: Nested) -> Iterator[Val def _fill_value_rules( configuration: RawBytesConfiguration, nested: Nested, value: RawBytesFillValue ) -> Iterator[ValidationProblem]: - """One byte value, an integer in `[0, 255]`, for each 8 of the size. + """One byte value for each 8 of the size. The spec's text says `N` values for `r`, but `N` counts bits (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L897-L900), @@ -89,7 +89,6 @@ def _fill_value_rules( yield ValidationProblem( (), f"expected {expected} byte values, got {len(value)}", "invalid_value" ) - yield from byte_value_problems(value) def _storage(configuration: RawBytesConfiguration, nested: Nested) -> StorageClass | None: diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/struct.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/struct.py index 3f3d284204..a5ded5ad8a 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/struct.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/struct.py @@ -10,14 +10,14 @@ from typing_extensions import ReadOnly, TypedDict from zarr_metadata._common import JSONValue -from zarr_metadata._json import ValidationProblem +from zarr_metadata._json import ValidationProblem, shown from zarr_metadata.v3._definition import ( DataTypeDefinition, DataTypeField, Nested, StorageClass, fill_value_problems, - named_configuration, + spelled_canonically, storage_of, ) @@ -96,10 +96,10 @@ def _rules(configuration: StructConfiguration, nested: Nested) -> Iterator[Valid ) field_type = nested.get(("fields", index, "data_type")) if field_type is not None and storage_of(field_type) == "variable_length": - written, _, _ = named_configuration(field_type.json) yield ValidationProblem( ("fields", index, "data_type"), - f"expected a data type of fixed size, got {written!r}, whose values vary in size", + f"expected a data type of fixed size, got {shown(field_type.name)}, whose values " + "vary in size", "invalid_value", ) @@ -115,6 +115,7 @@ def _fill_value_rules( not hold, leaves its fill value unjudged. """ names = [member["name"] for member in configuration["fields"]] + declared = set(names) for index, name in enumerate(names): if name not in value: yield ValidationProblem( @@ -125,10 +126,27 @@ def _fill_value_rules( if field_type is not None: yield from fill_value_problems(field_type, value[name], (name,)) for key in value: - if key not in names: + if key not in declared: yield ValidationProblem((key,), f"no struct field is named {key!r}", "unknown_key") +def _fill_value_canonical( + configuration: StructConfiguration, nested: Nested, value: StructFillValue +) -> dict[str, JSONValue]: + """Each field's fill value in the canonical spelling of that field's own type, in the order the fields are declared. + + A field type the scope did not read, or whose reading the struct's does + not hold, spells its fill value as written. + """ + spelled: dict[str, JSONValue] = {} + for index, member in enumerate(configuration["fields"]): + name = member["name"] + field_type = nested.get(("fields", index, "data_type")) + held = value[name] + spelled[name] = held if field_type is None else spelled_canonically(field_type, held) + return spelled + + def _storage(configuration: StructConfiguration, nested: Nested) -> StorageClass | None: """Its fields' values, packed together: numbers of several bytes if any field holds them, single bytes if every field is made of them. @@ -161,6 +179,7 @@ def _storage(configuration: StructConfiguration, nested: Nested) -> StorageClass rules=_rules, fill_value=StructFillValue, fill_value_rules=_fill_value_rules, + fill_value_canonical=_fill_value_canonical, storage=_storage, ) """The `struct` data type: a record of named fields, each field's type a nested field.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint16.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint16.py index 3578ff03fe..fa461df280 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint16.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint16.py @@ -4,10 +4,11 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/data-types/index.html """ -from typing import Final, Literal +from typing import Annotated, Final, Literal + +from annotated_types import Interval from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._integer import integer_fill_value_rules UINT16_DATA_TYPE_NAME: Final = "uint16" """The `data_type` value for the `uint16` type.""" @@ -15,7 +16,7 @@ Uint16DataTypeName = Literal["uint16"] """Literal type of the `data_type` field for `uint16`.""" -Uint16FillValue = int +Uint16FillValue = Annotated[int, Interval(ge=0, le=2**16 - 1)] """Permitted JSON shape of the `fill_value` field for `uint16`: a JSON integer in [0, 65535].""" @@ -23,7 +24,6 @@ name=UINT16_DATA_TYPE_NAME, configuration=EmptyConfiguration, fill_value=Uint16FillValue, - fill_value_rules=integer_fill_value_rules(0, 2**16 - 1), storage=multi_byte, ) """The `uint16` data type: a bare name, with nothing to configure; its fill value an integer in its range.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint32.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint32.py index 771d86804d..4cb844acec 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint32.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint32.py @@ -4,10 +4,11 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/data-types/index.html """ -from typing import Final, Literal +from typing import Annotated, Final, Literal + +from annotated_types import Interval from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._integer import integer_fill_value_rules UINT32_DATA_TYPE_NAME: Final = "uint32" """The `data_type` value for the `uint32` type.""" @@ -15,7 +16,7 @@ Uint32DataTypeName = Literal["uint32"] """Literal type of the `data_type` field for `uint32`.""" -Uint32FillValue = int +Uint32FillValue = Annotated[int, Interval(ge=0, le=2**32 - 1)] """Permitted JSON shape of the `fill_value` field for `uint32`: a JSON integer in [0, 2**32 - 1].""" @@ -23,7 +24,6 @@ name=UINT32_DATA_TYPE_NAME, configuration=EmptyConfiguration, fill_value=Uint32FillValue, - fill_value_rules=integer_fill_value_rules(0, 2**32 - 1), storage=multi_byte, ) """The `uint32` data type: a bare name, with nothing to configure; its fill value an integer in its range.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint64.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint64.py index 4e72073b64..82005090e2 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint64.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint64.py @@ -4,10 +4,11 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/data-types/index.html """ -from typing import Final, Literal +from typing import Annotated, Final, Literal + +from annotated_types import Interval from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, multi_byte -from zarr_metadata.v3.data_type._integer import integer_fill_value_rules UINT64_DATA_TYPE_NAME: Final = "uint64" """The `data_type` value for the `uint64` type.""" @@ -15,7 +16,7 @@ Uint64DataTypeName = Literal["uint64"] """Literal type of the `data_type` field for `uint64`.""" -Uint64FillValue = int +Uint64FillValue = Annotated[int, Interval(ge=0, le=2**64 - 1)] """Permitted JSON shape of the `fill_value` field for `uint64`: a JSON integer in [0, 2**64 - 1].""" @@ -23,7 +24,6 @@ name=UINT64_DATA_TYPE_NAME, configuration=EmptyConfiguration, fill_value=Uint64FillValue, - fill_value_rules=integer_fill_value_rules(0, 2**64 - 1), storage=multi_byte, ) """The `uint64` data type: a bare name, with nothing to configure; its fill value an integer in its range.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint8.py b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint8.py index 54ea7c59a7..46eb1ff93c 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint8.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/data_type/uint8.py @@ -4,10 +4,11 @@ See https://zarr-specs.readthedocs.io/en/latest/v3/data-types/index.html """ -from typing import Final, Literal +from typing import Annotated, Final, Literal + +from annotated_types import Interval from zarr_metadata.v3._definition import DataTypeDefinition, EmptyConfiguration, single_byte -from zarr_metadata.v3.data_type._integer import integer_fill_value_rules UINT8_DATA_TYPE_NAME: Final = "uint8" """The `data_type` value for the `uint8` type.""" @@ -15,7 +16,7 @@ Uint8DataTypeName = Literal["uint8"] """Literal type of the `data_type` field for `uint8`.""" -Uint8FillValue = int +Uint8FillValue = Annotated[int, Interval(ge=0, le=2**8 - 1)] """Permitted JSON shape of the `fill_value` field for `uint8`: a JSON integer in [0, 255].""" @@ -23,7 +24,6 @@ name=UINT8_DATA_TYPE_NAME, configuration=EmptyConfiguration, fill_value=Uint8FillValue, - fill_value_rules=integer_fill_value_rules(0, 2**8 - 1), storage=single_byte, ) """The `uint8` data type: a bare name, with nothing to configure; its fill value an integer in its range.""" diff --git a/packages/zarr-metadata/src/zarr_metadata/v3/definition.py b/packages/zarr-metadata/src/zarr_metadata/v3/definition.py index bdd28ba803..2a10ac806b 100644 --- a/packages/zarr-metadata/src/zarr_metadata/v3/definition.py +++ b/packages/zarr-metadata/src/zarr_metadata/v3/definition.py @@ -12,10 +12,12 @@ checked configuration has it as its static type. It reads as the typing spec defines a TypedDict -- `total`, `Required`, `NotRequired`, `closed` and `extra_items` mean what they mean to a type checker, - whether or not its module postpones annotations; + whether or not its module postpones annotations -- and a member's type + carries its bounds, as pydantic reads them: a gzip `level` is + `Annotated[int, Interval(ge=0, le=9)]`; - `rules`, a function over that TypedDict yielding what the spec - disallows -- a bound, members read together -- located in the - configuration. + disallows that a type cannot say -- members read together -- located + in the configuration. What kind of metadata a definition defines is its type: `CodecDefinition` (with the codec's `kind`, and its `size`: whether the size of what it gives @@ -26,43 +28,51 @@ **Reading JSON.** Three steps, each feeding the next, and each usable on its own by a caller that holds nothing but JSON: -1. `check(value, SomeTypedDict)`, from `zarr_metadata.typed_json` and - here too, type-checks JSON against a TypedDict and needs nothing else: +1. `check(value, SomeTypedDict)`, from `zarr_metadata.typed_json`, + type-checks JSON against a TypedDict and needs nothing else: a value of the TypedDict or None, and every problem, each located. The value holds what the TypedDict admits and nothing else. A member typed with a field alias is checked as the JSON a metadata field is. -2. `definition.judge(configuration)` is the check, each nested field's - envelope judged -- a stray member, a `must_understand` of `false` -- +2. `definition.read_configuration(configuration)` is the configuration as + that definition reads it: type-checked, each nested field's envelope judged -- a stray member, a `must_understand` of `false` -- and then the rules, for one configuration: - `GZIP_CODEC.judge({"level": 12})`. + `GZIP_CODEC.read_configuration({"level": 12})`. 3. `resolve(field, CodecDefinition, CORE_AND_EXTENSIONS)` reads a whole field in a scope: its envelope judged, its name related to a definition, its configuration judged, and each nested field read the - same way. It returns `Resolved` -- the field's JSON, its - `resolution`, the definition and the checked configuration -- and - every problem. A name nothing in scope claims is `out_of_scope`: - left unjudged, which is what keeps the format open. A field is read - as one of the five kinds, with or without type arguments; + same way. It returns what the scope made of the field, and every + problem: `AcceptedField` by the definition that claims its name, with the + configuration it checked and allowed and the fields it holds, each + read the same way; `UnclaimedField`, a name nothing in scope claims, left + unjudged, which is what keeps the format open; or `RefusedField`, whose + problems say why. Each has the field's JSON, the `name` it is + written with, the kind it was read as, `read_as`, the `definition` + that claims its name -- None for `UnclaimedField` -- and the fields it + holds as the scope read them, `nested`. The two a model holds, `AcceptedField` + and `UnclaimedField`, have a `configuration` and `to_json()`, the field as + a document writes it; two of them are equal when they read the same, + however each was spelled. `ResolvedField` is the three, for `match`. A + field is read as one of the five kinds, with or without type + arguments; `resolve(field, Definition, scope)` is a `TypeError`, since nothing - is filed under it. `configuration_of(resolved, GZIP_CODEC)` is the - configuration typed as that definition's TypedDict, when it read it. + is filed under it. A whole v3 array document is read by + `read_array_metadata_v3`, in `zarr_metadata.model`, whose reading + gives each field with where it sits in the document. from zarr_metadata.v3.codec.gzip import GZIP_CODEC - from zarr_metadata.v3.definition import ( - CORE_AND_EXTENSIONS, - CodecDefinition, - configuration_of, - resolve, - ) + from zarr_metadata.v3.definition import CORE_AND_EXTENSIONS, CodecDefinition, resolve resolved, problems = resolve({"name": "gzip", "configuration": {"level": 12}}, CodecDefinition, CORE_AND_EXTENSIONS) - resolved.resolution # 'invalid' + resolved # RefusedField(..., definition=CodecDefinition(name='gzip'), ...) problems[0].loc # ('configuration', 'level') + problems[0].input # 12 + dict(problems[0].ctx) # {'ge': 0, 'le': 9} resolved, problems = resolve({"name": "gzip", "configuration": {"level": 5}}, CodecDefinition, CORE_AND_EXTENSIONS) - configuration_of(resolved, GZIP_CODEC) # {'level': 5}, a GzipCodecConfiguration + resolved.definition is GZIP_CODEC # True + resolved.configuration # {'level': 5} Problems are values, not exceptions: `ValidationProblem(loc, message, kind)`, with `kind` one of `invalid_type`, `invalid_value`, @@ -70,19 +80,33 @@ unknown key is still read -- the key reported, the configuration judged without it -- so a consumer that tolerates one filters by kind and uses what was read; the field is valid only when there is no problem at all. +Each problem carries what its message says as data, as pydantic's errors +and zod's issues do: `input`, what was found at `loc` -- the `12` above -- +and `ctx`, what was expected, where that is more than a type: the bounds +`{"ge": 0, "le": 9}`, or the values of a closed set. -**Writing an extension.** A TypedDict, a function for its rules, and a -definition; then a scope that holds it. The TypedDict is a -`typing_extensions.TypedDict`: `closed` and `extra_items` are PEP 728's, -which `typing.TypedDict` does not take on the versions this package -supports. The rules are handed the configuration and the fields it holds -as the scope read them: a field that is read keeps what it read inside it -as `Resolved.nested`, a `Nested` mapping by where each sits, so a -struct's rules reach its field types. `judge`, which reads in no scope, -hands them none. +**Writing an extension.** A TypedDict, which says what the +configuration's JSON is, bounds and all; a function for the rules a type +cannot say; and a definition; then a scope that holds it. The TypedDict +is a `typing_extensions.TypedDict`: `closed` and `extra_items` are PEP +728's, which `typing.TypedDict` does not take on the versions this +package supports. The rules are handed the configuration and the fields +it holds as the scope read them: a field that is read keeps what it read +inside it as `AcceptedField.nested`, a `Nested` mapping by where each sits, so a +struct's rules reach its field types. `read_configuration`, which reads in no scope, +hands them none. A rule's message shows a value as the package's own +messages do, as JSON, with `shown`: `null`, `[1, 2]`, `"C"`. A rule +reports where a problem is; what is found there is the problem's +`input` without the rule saying so. Define each function at a module's +top level: a model holds the definitions that read its fields, so it +pickles, and compares equal once loaded, only when they do -- a lambda +or a closure does not pickle, and a `functools.partial` pickles but +compares unequal to itself loaded. from collections.abc import Iterator + from typing import Annotated, NotRequired + from annotated_types import Ge from typing_extensions import TypedDict from zarr_metadata.v3.definition import ( @@ -94,14 +118,18 @@ class AcmeLz4Configuration(TypedDict, closed=True): - acceleration: int + acceleration: Annotated[int, Ge(1)] + dictionary: NotRequired[str] + dictionary_size: NotRequired[Annotated[int, Ge(1)]] def acme_lz4_rules( configuration: AcmeLz4Configuration, nested: Nested ) -> Iterator[ValidationProblem]: - if configuration["acceleration"] < 1: - yield ValidationProblem(("acceleration",), "expected an integer >= 1", "invalid_value") + if "dictionary" in configuration and "dictionary_size" not in configuration: + yield ValidationProblem( + ("dictionary_size",), "a dictionary needs its size", "missing_key" + ) ACME_LZ4 = CodecDefinition( @@ -129,7 +157,20 @@ def acme_lz4_rules( word, so a definition refuses it. Its members are the shapes JSON takes: `int`, `float`, `bool`, `str`, `None`, `JSONValue`, a `Literal`, `tuple[T, ...]` and `tuple[T1, T2]`, a union, a TypedDict, -`Mapping[str, V]`, a `NewType` and a type alias. +`Mapping[str, V]`, a `NewType` and a type alias. A number's type may +carry bounds, as annotated-types spells them and pydantic reads them -- +`Gt`, `Ge`, `Lt`, `Le` and `Interval`, one from each side -- at any +depth: `tuple[Annotated[int, Ge(1)], ...]` bounds each element. A value +out of them is a problem, `invalid_value`, whose message says what the +type admits, "expected an integer >= 1, got 0", and whose `ctx` holds +the bounds. The rules are asked only of a configuration within its +bounds, so a rule relies on them, as pydantic's after-validators and +zod's refinements do: until a value out of bounds is fixed, it is the +one problem reported of the configuration. `Annotated` may also carry a +note, a string or a `Doc`. Any other metadata -- a `MinLen`, a +`Predicate`, pydantic's `Field` -- is refused when the definition is +built, since a type the checker does not hold its values to would say +what is not so. A member holding another metadata field is annotated with the field alias of its kind -- a shard's `codecs: tuple[CodecField, ...]` -- and read in @@ -143,18 +184,28 @@ def acme_lz4_rules( unjudged, its size unknown with the rest of it. `ZarrV3MetadataFieldJSON` is the same JSON, but checks as JSON and nothing more, so a definition refuses a member typed with it. An extension with nothing to configure -takes `EmptyConfiguration`, and is written as its bare name. +takes `EmptyConfiguration`, and is written with its name alone. A data type also says what its fill value is: `fill_value`, the JSON -shape of one as an annotation the checker reads -- `Int8FillValue` -- and -`fill_value_rules`, a function yielding what the spec disallows in a fill -value of that shape: an integer out of range, a hex string of another -width. The rules are handed the configuration, the fields it holds as the +shape of one as an annotation the checker reads -- `Int8FillValue`, +whose type carries the range -- and `fill_value_rules`, a function +yielding what the spec disallows in a fill value of that shape that the +type cannot say: a hex string of another width, a number of byte values +the size does not take. The rules are handed the configuration, the fields it holds as the scope read them, and the typed fill value, so a struct judges each field's fill value by that field's own type. `fill_value_problems(data_type, value)` judges a fill value against a data type field the scope read; one nothing in scope claims leaves it unjudged. A data type that says nothing of its fill value takes any JSON. +A fill value may be spelled more ways than one -- `"NaN"` and +`"0x7fc00000"` are one `float32` -- so a data type says which spelling +is its value's own: `fill_value_canonical`, handed what the rules are +handed and a fill value they allow. Two fill values are one value of the +type exactly when their canonical spellings are written alike: the same +JSON, as `json.dumps` writes it, which `==` is not -- it takes `-0.0`, +a `float32` of its own, for `0.0`. `canonical_fill_value(data_type, +value)` spells one, and gives `UNSET` for a fill value with a problem; a +data type that says nothing of it spells each value as written. A data type says how its values are stored, too: `storage`, a function of its configuration and the fields it holds, giving a `StorageClass` -- @@ -170,9 +221,9 @@ def acme_lz4_rules( fall short of one -- located in the configuration. A grid that says nothing of the shape fits every one. It also says the lengths its chunks take along each axis of an array it fits, `chunk_lengths`: a set per -axis, since a rectilinear grid's chunks differ. `chunk_grid_lengths(grid, -shape)` gives both of a chunk grid field the scope read: an entry for -each dimension of the shape, None where nothing says the lengths. +axis, since a rectilinear grid's chunks differ. A reading holds both of +the grid it read: an entry for each dimension of the shape, None where +nothing says the lengths. A codec is judged against what it is handed. The array hands its first codec a `Chunk`: the lengths of its grid's chunks along each of the @@ -181,11 +232,10 @@ def acme_lz4_rules( a chunk: `chunk_rules`, located in its configuration -- a `transpose` whose `order` has another number of axes. An array -> array codec says what it hands the next, whatever its chunk rules found: `transition` -- -`transpose` permutes the axes. `read_pipeline(codecs, chunk)` reads codec -fields the scope read as a pipeline: their order -- array -> array -codecs, one array -> bytes codec, bytes -> bytes codecs -- and then each -against the chunk it is handed, giving each codec's `Stage` with that -chunk. A codec that holds pipelines of its own says what each is +`transpose` permutes the axes. A reading reads the codec fields as a +pipeline: their order -- array -> array codecs, one array -> bytes codec, +bytes -> bytes codecs -- and then each against the chunk it is handed, +giving each codec's `Stage` with that chunk. A codec that holds pipelines of its own says what each is handed: `pipelines`, by the member of its configuration that holds each -- a shard's inner codecs its inner chunks, its index codecs the shard index -- and each is read the same way, its stages kept as the codec's @@ -202,16 +252,48 @@ def acme_lz4_rules( refused, since that name reads as `r*`. `r*` itself is notation, and a document that writes it names nothing in any scope. -**The simplest spelling.** `canonicalize(field, kind, scope)` gives a -field without problems in its simplest equivalent spelling: each nested +**The simplest spelling.** A field without problems has a simplest +equivalent spelling, which is what two fields are compared by: each nested field in its own simplest spelling, then the definition's `canonical` -- blosc drops a `typesize` that `noshuffle` ignores, a rectilinear grid run-length encodes its chunk shapes -- and the envelope in the fewest -words; raw bits write their size back into the name, in decimal, so -`r008` is `r8`. A field with any problem, an unknown key included, has -none: a simpler spelling of it would erase what its author wrote. What -`canonical` gives is judged again: one that does not hold is a -`ValueError`, a fault in the definition. +words every reader takes: a data type with nothing to configure is its +bare name, any other field an object, `{"name": ...}`, as a Zarr v3.0 +reader takes no short-hand name in `codecs`; raw bits write their size +back into the name, in decimal, so `r008` is `r8`. A field with any +problem, an unknown key included, has none: a simpler spelling of it +would erase what its author wrote. What `canonical` gives is judged +again: one that does not hold is a `ValueError`, a fault in the +definition. `canonical_fill_value` spells a fill value the same way, as +the data type that read it spells one. + +**JSON Schema.** `node_metadata_json_schema_v3`, in `zarr_metadata.model`, +writes a whole `zarr.json` as a JSON Schema, draft 2020-12, for a +validator in another language or an editor, its fields as the scope +reads them: each definition's field -- +its name, its configuration as its TypedDict says, bounds and all, a +`must_understand` of `true`, and its bare name when it needs no +configuration -- and a name nothing in scope claims, with any +configuration. A field a configuration holds is written in the same +scope, and a member taking codecs of static size only takes those. The +rules are not in it, so a field it accepts may still have a problem; +one `resolve` reads without a problem, it accepts, as JSON: arrays as +lists, as a parser gives them. Each configuration +TypedDict, and each field alias, is written once, in `$defs`, under its +name; the fill value is held to its data type's. + +**Scopes as values.** A scope reads documents of one Zarr format, the +`format` every kind it files declares -- `ZarrV3Context` is the type of +a v3 scope, `ZarrV2Context` of a v2 one -- and a reader refuses a scope +of the other format with `TypeError`, while a scope that files nothing +is of no format and reads in either. Two scopes are equal when they file +the same definitions, and equal scopes hash alike. `Context.joined(*scopes)` is +the least scope above each, or a `ScopeConflictError` naming each name +filed two ways; `extended_with` remains the way to take a name over on +purpose. A model's `refined_in` moves it to a scope that claims more and +contradicts nothing, and `refines` orders two models by information: a +name nothing claimed, read by a definition, is a gain; the reverse a +loss; one name read by two definitions a conflict. A definition checks itself when it is built, and each of these is a `TypeError` saying what is wrong: a `configuration` that is not a @@ -220,68 +302,82 @@ def acme_lz4_rules( that is not a string; a member declared as a function -- `rules`, `canonical`, `fill_value_rules`, `storage`, `shape_rules`, `chunk_lengths`, `chunk_rules`, `transition`, `pipelines` -- that is not -one; a data type's `fill_value` no checker reads; a codec `kind` that is -not one of the three, or a `size` that is not `"static"` or `"dynamic"`; -a function no codec of its kind is asked -- chunk rules or pipelines of -a bytes -> bytes codec, which is handed bytes, or a `transition` of a -codec that hands on bytes; a data type named as raw bits of one size are -written. A scope refuses a definition of no kind. Nothing happens at -class creation. +one; a data type's `fill_value` no checker reads, or one holding a +metadata field, which a value of the data type never is; a codec `kind` +that is not one of the three, or a `size` that is not `"static"` or +`"dynamic"`; a function no codec of its kind is asked -- chunk rules or +pipelines of a bytes -> bytes codec, which is handed bytes, or a +`transition` of a codec that hands on bytes; a data type named as raw +bits of one size are written. A scope refuses a definition of no kind, +and a kind is declared with its format, `class Tag(Definition[C], +kind=True, format=3)`. Nothing else happens at class creation. """ from zarr_metadata._common import JSONValue -from zarr_metadata._json import MetadataValidationError, ProblemKind, ValidationProblem -from zarr_metadata._typed_json import Loc, check -from zarr_metadata.v3._common import ZarrV3MetadataFieldJSON +from zarr_metadata._json import MetadataValidationError, ProblemKind, ValidationProblem, shown +from zarr_metadata._typed_json import Loc +from zarr_metadata.v3._common import ( + ChunkGridField, + ChunkKeyEncodingField, + CodecField, + DataTypeField, + StaticCodecField, + StorageTransformerField, +) from zarr_metadata.v3._definition import ( + AcceptedField, Chunk, ChunkGridDefinition, - ChunkGridField, ChunkKeyEncodingDefinition, - ChunkKeyEncodingField, CodecDefinition, - CodecField, CodecKind, CodecSize, DataTypeDefinition, - DataTypeField, Definition, EmptyConfiguration, Lengths, Nested, - Resolution, - Resolved, - StaticCodecField, + RefusedField, + ResolvedField, StorageClass, StorageTransformerDefinition, - StorageTransformerField, - Unread, - canonicalize, - chunk_grid_lengths, - configuration_of, + UnclaimedField, + canonical_fill_value, fill_value_problems, resolve, storage_of, ) -from zarr_metadata.v3._pipeline import Stage, read_pipeline -from zarr_metadata.v3._registry import CORE, CORE_AND_EXTENSIONS, Context +from zarr_metadata.v3._pipeline import Stage +from zarr_metadata.v3._registry import CORE, CORE_AND_EXTENSIONS, Context, ZarrV3Context +from zarr_metadata.v3._scope import ( + ClaimKey, + Claims, + Conflict, + Disagreements, + ScopeConflictError, +) __all__ = [ "CORE", "CORE_AND_EXTENSIONS", + "AcceptedField", "Chunk", "ChunkGridDefinition", "ChunkGridField", "ChunkKeyEncodingDefinition", "ChunkKeyEncodingField", + "ClaimKey", + "Claims", "CodecDefinition", "CodecField", "CodecKind", "CodecSize", + "Conflict", "Context", "DataTypeDefinition", "DataTypeField", "Definition", + "Disagreements", "EmptyConfiguration", "JSONValue", "Lengths", @@ -289,22 +385,20 @@ class creation. "MetadataValidationError", "Nested", "ProblemKind", - "Resolution", - "Resolved", + "RefusedField", + "ResolvedField", + "ScopeConflictError", "Stage", "StaticCodecField", "StorageClass", "StorageTransformerDefinition", "StorageTransformerField", - "Unread", + "UnclaimedField", "ValidationProblem", - "ZarrV3MetadataFieldJSON", - "canonicalize", - "check", - "chunk_grid_lengths", - "configuration_of", + "ZarrV3Context", + "canonical_fill_value", "fill_value_problems", - "read_pipeline", "resolve", + "shown", "storage_of", ] diff --git a/packages/zarr-metadata/tests/model/conftest.py b/packages/zarr-metadata/tests/model/conftest.py new file mode 100644 index 0000000000..ea7c08e22a --- /dev/null +++ b/packages/zarr-metadata/tests/model/conftest.py @@ -0,0 +1,22 @@ +"""Fixtures for the model tests.""" + +from __future__ import annotations + +import sys +from typing import TYPE_CHECKING + +import pytest + +if TYPE_CHECKING: + from collections.abc import Iterator + + +@pytest.fixture +def interpreter_writes_4300_digits() -> Iterator[None]: + """The interpreter's default limit on writing an integer, pinned: an environment may lift it (`PYTHONINTMAXSTRDIGITS=0`), and a test of what happens past it needs it.""" + limit = sys.get_int_max_str_digits() + sys.set_int_max_str_digits(4300) + try: + yield + finally: + sys.set_int_max_str_digits(limit) diff --git a/packages/zarr-metadata/tests/model/test_array.py b/packages/zarr-metadata/tests/model/test_array.py index 768a9c0d0c..f8b4590bd6 100644 --- a/packages/zarr-metadata/tests/model/test_array.py +++ b/packages/zarr-metadata/tests/model/test_array.py @@ -7,44 +7,63 @@ import pickle from collections import UserDict from collections.abc import Callable -from typing import TYPE_CHECKING, TypeGuard, get_args, get_origin, get_type_hints +from typing import TYPE_CHECKING, Any, TypeGuard, cast, get_args, get_origin, get_type_hints import pytest from typing_extensions import Unpack from tests.model._cases import Expect, ExpectFail, mutate_nested_containers -from zarr_metadata._json import arrays_to_tuples, prefixed +from zarr_metadata._json import ( + JSON_DEPTH, + arrays_to_tuples, + is_json, + json_text, + parse_json, + prefixed, + validate_json, +) from zarr_metadata.model import ( - ARRAY_METADATA_OPTIONAL_KEYS_V3, - ARRAY_METADATA_REQUIRED_KEYS_V3, - ARRAY_METADATA_STANDARD_KEYS_V3, UNSET, MetadataValidationError, ValidationProblem, ZarrV2ArrayMetadata, - ZarrV2ArrayMetadataPartial, + ZarrV2ArrayMetadataUpdate, ZarrV3ArrayMetadata, - ZarrV3ArrayMetadataPartial, - ZarrV3MetadataField, - ZarrV3NamedConfig, + ZarrV3GroupMetadata, is_array_metadata_v2, is_array_metadata_v3, is_group_metadata_v2, is_group_metadata_v3, - is_json, is_metadata_field_v3, parse_array_metadata_v2, parse_array_metadata_v3, - parse_json, parse_metadata_field_v3, validate_array_metadata_v2, validate_array_metadata_v3, - validate_json, validate_metadata_field_v3, ) +from zarr_metadata.model._validation import ( + ARRAY_METADATA_OPTIONAL_KEYS_V3, + ARRAY_METADATA_REQUIRED_KEYS_V3, + ARRAY_METADATA_STANDARD_KEYS_V3, +) +from zarr_metadata.v3._definition import ( + canonical_of, + configuration_of, + fields_of, + with_problems, +) +from zarr_metadata.v3.array import ZarrV3ArrayMetadataJSONPartial +from zarr_metadata.v3.codec.gzip import GZIP_CODEC +from zarr_metadata.v3.definition import ( + CORE, + CORE_AND_EXTENSIONS, + AcceptedField, + UnclaimedField, +) if TYPE_CHECKING: - from zarr_metadata._common import JSONValue + from zarr_metadata._common import JSONValue, ZarrV3NamedConfigJSON from zarr_metadata.v2 import ZarrV2CodecMetadata # --- public exports -------------------------------------------------------- @@ -55,8 +74,6 @@ def test_guards_exported_from_package() -> None: import zarr_metadata.model for name in ( - "is_json", - "parse_json", "is_metadata_field_v3", "parse_metadata_field_v3", "is_array_metadata_v3", @@ -140,7 +157,6 @@ def test_validation_diagnostics_exported_from_package() -> None: for name in ( "ValidationProblem", "MetadataValidationError", - "validate_json", "validate_metadata_field_v3", "validate_array_metadata_v3", "validate_array_metadata_v2", @@ -188,81 +204,18 @@ def test_string_nan_fill_value_roundtrips() -> None: # The string form round-trips cleanly under default dataclass equality, # unlike a raw float('nan'), which is not JSON. """A float array's string 'NaN' fill_value round-trips cleanly.""" - m = ZarrV3ArrayMetadata.create_default( - fill_value="NaN", - data_type=ZarrV3NamedConfig(name="float32", configuration={}), - codecs=(ZarrV3NamedConfig(name="bytes", configuration={"endian": "little"}),), - ) + m = ZarrV3ArrayMetadata.create_default(fill_value="NaN", data_type="float32") assert ZarrV3ArrayMetadata.from_json(m.to_json()) == m assert ZarrV3ArrayMetadata.from_json(m.to_json()).fill_value == "NaN" -# --- ZarrV3NamedConfig.to_json ------------------------------------------------ - -ZARR_TO_JSON_CASES = [ - Expect( - ZarrV3NamedConfig(name="regular", configuration={"chunk_shape": [1]}), - {"name": "regular", "configuration": {"chunk_shape": [1]}}, - id="with-configuration", - ), - Expect( - ZarrV3NamedConfig(name="bytes", configuration={}), - "bytes", - id="empty-configuration-shorthand", - ), -] - - -@pytest.mark.parametrize("case", ZARR_TO_JSON_CASES, ids=lambda c: c.id) -def test_zarr_metadata_v3_to_json(case: Expect[ZarrV3NamedConfig, object]) -> None: - """ZarrV3NamedConfig.to_json emits the canonical extension form.""" - assert case.input.to_json() == case.output - - -def test_zarr_metadata_v3_to_json_preserves_false_obligation() -> None: - """An empty optional extension stays an object so false is not lost.""" - model = ZarrV3NamedConfig(name="optional", configuration={}, must_understand=False) - assert model.to_json() == {"name": "optional", "must_understand": False} - - -# --- ZarrV3NamedConfig.from_json ----------------------------------------------- - -ZARR_FROM_JSON_CASES = [ - Expect("bytes", ZarrV3NamedConfig(name="bytes", configuration={}), id="bare-string"), - Expect( - {"name": "regular", "configuration": {"chunk_shape": [1]}}, - ZarrV3NamedConfig(name="regular", configuration={"chunk_shape": (1,)}), - id="object-with-config", - ), - Expect( - {"name": "bytes"}, - ZarrV3NamedConfig(name="bytes", configuration={}), - id="object-without-config", - ), -] - - -@pytest.mark.parametrize("case", ZARR_FROM_JSON_CASES, ids=lambda c: c.id) -def test_zarr_metadata_v3_from_json(case: Expect[object, ZarrV3NamedConfig]) -> None: - """ZarrV3NamedConfig.from_json parses both the bare-string and object forms.""" - assert ZarrV3NamedConfig.from_json(case.input) == case.output - - -def test_zarr_metadata_v3_from_json_preserves_false_obligation() -> None: - """Explicit false is represented on the normalized model.""" - model = ZarrV3NamedConfig.from_json({"name": "optional", "must_understand": False}) - assert model.must_understand is False - - # --- V3 baseline ----------------------------------------------------------- def test_v3_to_json_emits_canonical_document() -> None: """V3 to_json emits exactly the expected document (which covers every spec-required key by construction).""" - out = ZarrV3ArrayMetadata.create_default( - shape=(10,), data_type=ZarrV3NamedConfig(name="int32", configuration={}) - ).to_json() + out = ZarrV3ArrayMetadata.create_default(shape=(10,), data_type="int32").to_json() assert out == { "zarr_format": 3, "node_type": "array", @@ -270,23 +223,46 @@ def test_v3_to_json_emits_canonical_document() -> None: "fill_value": 0, "data_type": "int32", "chunk_grid": {"name": "regular", "configuration": {"chunk_shape": (10,)}}, - "codecs": ("bytes",), - "chunk_key_encoding": "default", + "codecs": ({"name": "bytes", "configuration": {"endian": "little"}},), + "chunk_key_encoding": {"name": "default"}, } +@pytest.mark.parametrize( + "data_type", + ["int32", {"name": "int32"}, {"name": "int32", "configuration": {}}], + ids=["bare-name", "object", "empty-configuration"], +) +def test_v3_a_data_type_with_nothing_to_configure_is_written_by_its_bare_name( + data_type: object, +) -> None: + # As core data types have been written since Zarr v3.0, which is how + # zarr-python reads them; every other extension point is an object. + document = { + **ZarrV3ArrayMetadata.create_default(shape=(10,)).to_json(), + "data_type": data_type, + "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}], + } + model = ZarrV3ArrayMetadata.from_json(document) + # The document is written as it was written; the field alone, by its + # bare name, which every reader takes. + assert model.to_json()["data_type"] == data_type + assert model.data_type.to_json() == "int32" + + def test_v3_dimension_names_included_when_present() -> None: """V3 to_json includes dimension_names when they are set.""" out: dict[str, object] = dict( - ZarrV3ArrayMetadata.create_default(dimension_names=("x",)).to_json() + ZarrV3ArrayMetadata.create_default(shape=(4,), dimension_names=("x",)).to_json() ) assert out["dimension_names"] == ("x",) def test_v3_dimension_names_omitted_when_none() -> None: """V3 to_json omits dimension_names when they are UNSET.""" - out = ZarrV3ArrayMetadata.create_default(dimension_names=UNSET).to_json() - assert "dimension_names" not in out + model = ZarrV3ArrayMetadata.create_default() + assert model.dimension_names is UNSET + assert "dimension_names" not in model.to_json() # --- BUG 1: attributes gated on dimension_names ---------------------------- @@ -299,11 +275,9 @@ def test_v3_attributes_included_when_dimension_names_is_none() -> None: so non-empty attributes were silently dropped when there were no dimension names. """ - out: dict[str, object] = dict( - ZarrV3ArrayMetadata.create_default( - dimension_names=UNSET, attributes={"foo": "bar"} - ).to_json() - ) + model = ZarrV3ArrayMetadata.create_default(attributes={"foo": "bar"}) + assert model.dimension_names is UNSET + out: dict[str, object] = dict(model.to_json()) assert out["attributes"] == {"foo": "bar"} @@ -316,17 +290,18 @@ def test_v3_single_storage_transformer_included() -> None: Regression: the guard used ``> 1`` instead of ``> 0``, dropping a lone storage transformer. """ - st = ZarrV3NamedConfig(name="some_transformer", configuration={}) + st: ZarrV3NamedConfigJSON = {"name": "some_transformer"} out: dict[str, object] = dict( ZarrV3ArrayMetadata.create_default(storage_transformers=(st,)).to_json() ) - assert out["storage_transformers"] == ("some_transformer",) + assert out["storage_transformers"] == ({"name": "some_transformer"},) def test_v3_no_storage_transformers_omitted() -> None: - """V3 to_json omits storage_transformers when empty.""" - out = ZarrV3ArrayMetadata.create_default(storage_transformers=()).to_json() - assert "storage_transformers" not in out + """V3 to_json omits storage_transformers when the document wrote none, and writes an empty one as written.""" + assert "storage_transformers" not in ZarrV3ArrayMetadata.create_default().to_json() + written = dict(ZarrV3ArrayMetadata.create_default(storage_transformers=()).to_json()) + assert written["storage_transformers"] == () # --- V3 extra fields ------------------------------------------------------- @@ -334,16 +309,9 @@ def test_v3_no_storage_transformers_omitted() -> None: def test_v3_extra_fields_merged() -> None: """V3 to_json merges extra_fields into the top-level document.""" - out = ZarrV3ArrayMetadata.create_default( - extra_fields={"my_ext": {"must_understand": False}} - ).to_json() - assert out["my_ext"] == {"must_understand": False} - - -def test_v3_extra_fields_overlapping_standard_field_rejected() -> None: - """Constructing a V3 model with an extra field that collides with a standard key is rejected.""" - with pytest.raises(ValueError): - ZarrV3ArrayMetadata.create_default(extra_fields={"shape": {"must_understand": False}}) + model = ZarrV3ArrayMetadata.create_default(my_ext={"must_understand": False}) + assert model.extra_fields == {"my_ext": {"must_understand": False}} + assert model.to_json()["my_ext"] == {"must_understand": False} # --- V3 key/value ---------------------------------------------------------- @@ -387,7 +355,7 @@ def test_v3_create_default_is_valid_empty_array() -> None: """V3 create_default builds a structurally valid empty array that round-trips.""" m = ZarrV3ArrayMetadata.create_default() assert m.shape == () - assert m.data_type == ZarrV3NamedConfig(name="uint8", configuration={}) + assert m.data_type.to_json() == "uint8" assert m.fill_value == 0 assert m.attributes == {} assert m.extra_fields == {} @@ -402,7 +370,7 @@ def test_v3_create_default_applies_overrides() -> None: assert m.shape == (4, 4) assert m.attributes == {"a": 1} # un-overridden fields keep their defaults - assert m.data_type == ZarrV3NamedConfig(name="uint8", configuration={}) + assert m.data_type.to_json() == "uint8" def test_v2_create_default_is_valid_empty_array() -> None: @@ -423,7 +391,7 @@ def test_v2_create_default_applies_overrides() -> None: m = ZarrV2ArrayMetadata.create_default(shape=(8,), attributes={"k": "v"}) assert m.shape == (8,) assert m.attributes == {"k": "v"} - assert m.dtype == "|u1" # default dtype unchanged + assert m.dtype.to_json() == "|u1" # default dtype unchanged # --- V3 update ------------------------------------------------------------- @@ -460,34 +428,255 @@ def test_update_no_args_returns_equal_model( ) -> None: """update with no arguments returns a model equal to the original.""" base = model_cls.create_default() - assert base.update() == base + updated = base.update() + assert updated == base # V3-only update tests — kept direct (extra_fields is v3-specific) -def test_update_can_replace_extra_fields() -> None: - """update can replace the extra_fields mapping.""" - base = ZarrV3ArrayMetadata.create_default(extra_fields={}) - updated = base.update(extra_fields={"my_ext": {"must_understand": False}}) +def test_update_can_add_an_extension_member() -> None: + """update can add a member the spec does not define, which the model holds in extra_fields.""" + base = ZarrV3ArrayMetadata.create_default() + updated = base.update(my_ext={"must_understand": False}) assert updated.extra_fields == {"my_ext": {"must_understand": False}} -def test_update_replaces_extra_fields_rather_than_merging() -> None: - """update replaces extra_fields wholesale rather than merging.""" - base = ZarrV3ArrayMetadata.create_default(extra_fields={"a": {"must_understand": False}}) - updated = base.update(extra_fields={"b": {"must_understand": True}}) - assert updated.extra_fields == {"b": {"must_understand": True}} +def test_update_reads_every_member_in_the_models_own_scope() -> None: + """`update` reads the document it makes in the scope the model was read in, so a name that scope leaves unclaimed stays unclaimed; `with_context` is how another scope reads it. + `zstd` is an extension, which `CORE` leaves unclaimed and unjudged. + """ + little: ZarrV3NamedConfigJSON = {"name": "bytes", "configuration": {"endian": "little"}} + zstd: ZarrV3NamedConfigJSON = {"name": "zstd", "configuration": {"level": 3, "checksum": False}} + base = ZarrV3ArrayMetadata.create_default(context=CORE, codecs=(little, zstd)) + assert isinstance(base.codecs[1], UnclaimedField) + kept = base.update(attributes={"k": 1}) + assert kept.codecs[1] == base.codecs[1] + assert kept.context == CORE + given = base.with_context(CORE_AND_EXTENSIONS).update(codecs=(little, zstd)) + assert isinstance(given.codecs[1], AcceptedField) + + +def test_update_leaves_out_a_member_given_as_unset() -> None: + """Each member a document may leave out: `dimension_names`, `attributes`, one the spec does not define.""" + base = ZarrV3ArrayMetadata.create_default( + shape=(2,), dimension_names=("x",), attributes={"a": 1}, my_ext={"must_understand": False} + ) + for updated, member in ( + (base.update(dimension_names=UNSET), "dimension_names"), + (base.update(attributes=UNSET), "attributes"), + (base.update(my_ext=UNSET), "my_ext"), + ): + written = updated.to_json() + assert member not in written + assert {key: value for key, value in base.to_json().items() if key != member} == written -def test_partial_keys_match_settable_model_fields() -> None: - """The partial TypedDict must list exactly the constructor-settable fields. - Guards against drift: adding/removing a settable field on the model - without updating ``ZarrV3ArrayMetadataPartial`` fails here. - """ - settable = {f.name for f in dataclasses.fields(ZarrV3ArrayMetadata) if f.init} - assert set(ZarrV3ArrayMetadataPartial.__annotations__) == settable +@pytest.mark.parametrize( + "members", + [ + {"data_type": {"name": "uint8"}}, + {"codecs": ("bytes",)}, + {"codecs": ({"name": "bytes", "configuration": {}, "must_understand": True},)}, + {"chunk_key_encoding": {"name": "default", "configuration": {}}}, + {"codecs": ({"name": "bytes"}, {"name": "acme.codec", "configuration": {}})}, + { + "shape": (4,), + "codecs": ( + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": (2,), + "codecs": ({"name": "bytes"},), + "index_codecs": ( + {"name": "bytes", "configuration": {"endian": "little"}}, + "crc32c", + ), + }, + }, + ), + }, + ], + ids=[ + "data-type-object", + "codec-bare", + "codec-verbose", + "empty-configuration", + "unclaimed", + "a-field-a-codec-holds-bare", + ], +) +def test_a_model_is_what_its_document_says_not_how_it_is_spelled( + members: ZarrV3ArrayMetadataJSONPartial, +) -> None: + """A model reads back from its own document as itself, and equals the model of any spelling of it.""" + model = ZarrV3ArrayMetadata.create_default(**members) + assert ZarrV3ArrayMetadata.from_json(model.to_json()) == model + written = {**model.to_json(), **members} + assert ZarrV3ArrayMetadata.from_json(written) == model + + +_DATETIME = {"name": "numpy.datetime64", "configuration": {"unit": "s", "scale_factor": 1}} +_LITTLE = {"name": "bytes", "configuration": {"endian": "little"}} + + +@pytest.mark.parametrize( + ("data_type", "left", "right", "same"), + [ + ("float32", "NaN", "0x7fc00000", True), + ("float32", 1, 1.0, True), + ("float32", 0.1, 0.10000000149011612, True), + ("float32", 0.0, -0.0, False), + ("float32", "NaN", "0xffc00000", False), + (_DATETIME, "NaT", -(2**63), True), + ("bytes", [65], "QR==", True), + # The fill value of a data type nothing in scope claims is not + # interpreted, and is compared as JSON text. + ("acme.decimal", {"a": 1, "b": 2}, {"b": 2, "a": 1}, True), + ("acme.decimal", [1, 2], [2, 1], False), + ("acme.decimal", True, 1, False), + ("acme.decimal", 0.0, -0.0, False), + ], +) +def test_two_models_are_one_array_when_their_fill_values_are_one_value( + data_type: object, left: object, right: object, same: bool +) -> None: + """As two fields are one when they read the same, however each is spelled.""" + codecs = [{"name": "vlen-bytes"} if data_type == "bytes" else _LITTLE] + document = {**ZarrV3ArrayMetadata.create_default(shape=(2,)).to_json(), "codecs": codecs} + models = [ + ZarrV3ArrayMetadata.from_json({**document, "data_type": data_type, "fill_value": value}) + for value in (left, right) + ] + assert (models[0] == models[1]) is same + assert (models[1] == models[0]) is same + assert models[0] == ZarrV3ArrayMetadata.from_json(models[0].to_json()) + # Equal models hash alike. + if same: + assert hash(models[0]) == hash(models[1]) + # A group holding the arrays says so too. + groups = [ + ZarrV3GroupMetadata.create_default( + consolidated_metadata={**_INLINE, "metadata": {"a": model.to_json()}} + ) + for model in models + ] + assert (groups[0] == groups[1]) is same + + +_INLINE: dict[str, Any] = {"kind": "inline", "must_understand": False} +_NOSHUFFLE = {"cname": "lz4", "clevel": 5, "shuffle": "noshuffle", "blocksize": 0} + + +@pytest.mark.parametrize( + ("member", "left", "right"), + [ + ( + "codecs", + [_LITTLE, {"name": "blosc", "configuration": _NOSHUFFLE}], + [_LITTLE, {"name": "blosc", "configuration": {**_NOSHUFFLE, "typesize": 4}}], + ), + ( + "codecs", + [_LITTLE, {"name": "zstd", "configuration": {"level": 1}}], + [_LITTLE, {"name": "zstd", "configuration": {"level": 1, "checksum": False}}], + ), + ( + "chunk_key_encoding", + "default", + {"name": "default", "configuration": {"separator": "/"}}, + ), + ], + ids=["blosc-typesize-noshuffle", "zstd-checksum-false", "default-separator"], +) +def test_two_models_are_one_array_when_their_fields_read_the_same( + member: str, left: object, right: object +) -> None: + """The spec's equivalences, which each definition's `canonical` folds, hold in a model's `==`, though `to_json` writes each as given.""" + document = ZarrV3ArrayMetadata.create_default(shape=(2,)).to_json() + models = [ZarrV3ArrayMetadata.from_json({**document, member: value}) for value in (left, right)] + assert models[0] == models[1] + assert hash(models[0]) == hash(models[1]) + assert models[0].to_json() != models[1].to_json() + + +@pytest.mark.parametrize( + ("left", "right", "same"), + [ + ({"a": math.nan}, {"a": math.nan}, True), + ({"a": 1, "b": 2}, {"b": 2, "a": 1}, True), + ({"a": True}, {"a": 1}, False), + ({"a": 0.0}, {"a": -0.0}, False), + ({"a": 1}, {"a": 1.0}, False), + ], + ids=["nan", "key-order", "bool-vs-int", "signed-zero", "int-vs-float"], +) +def test_user_json_compares_as_text( + left: "dict[str, JSONValue]", right: "dict[str, JSONValue]", same: bool +) -> None: + """Attributes, which nothing interprets, compare as a document writes them: `NaN` is itself, `true` is not `1`, `-0.0` is not `0.0`.""" + arrays = [ZarrV3ArrayMetadata.create_default(attributes=held) for held in (left, right)] + groups = [ZarrV3GroupMetadata.create_default(attributes=held) for held in (left, right)] + v2 = [ZarrV2ArrayMetadata.create_default(attributes=held) for held in (left, right)] + for models in (arrays, groups, v2): + assert (models[0] == models[1]) is same + if same: + assert hash(models[0]) == hash(models[1]) + + +def test_a_model_holding_nan_user_data_equals_its_copies() -> None: + # What Python's `==` on the value denies: `nan != nan`. + model = ZarrV3ArrayMetadata.create_default(attributes={"_FillValue": math.nan}) + assert ZarrV3ArrayMetadata.from_key_value(model.to_key_value()) == model + assert pickle.loads(pickle.dumps(model)) == model + assert ZarrV3ArrayMetadata.from_json(json.loads(json.dumps(model.to_json()))) == model + + +@pytest.mark.parametrize( + ("left", "right", "same"), + [ + (0.0, 0.0, True), + ("NaN", "NaN", True), + (0.0, -0.0, False), + (1, 1.0, True), + ("Infinity", "NaN", False), + ], +) +def test_two_v2_models_are_one_array_when_their_documents_are_written_alike( + left: object, right: object, same: bool +) -> None: + """A v2 model compares by what its document means: a fill value an integer or a float is one value of a float type, `0.0` and `-0.0` two, and `NaN` is itself.""" + model = ZarrV2ArrayMetadata.create_default(shape=(2,), chunks=(2,), dtype=" None: + model = ZarrV3ArrayMetadata.create_default( + codecs=( + {"name": "bytes", "configuration": {"endian": "little"}}, + {"name": "gzip", "configuration": {"level": 1}}, + ), + context=CORE, + ) + for again in (pickle.loads(pickle.dumps(model)), copy.copy(model), copy.deepcopy(model)): + assert again == model + assert configuration_of(again.codecs[1], GZIP_CODEC) == {"level": 1} + + +def test_update_replaces_a_member_rather_than_merging_into_it() -> None: + """update replaces each member it is given whole, and keeps the others.""" + base = ZarrV3ArrayMetadata.create_default( + a={"must_understand": False, "x": 1}, b={"must_understand": False} + ) + updated = base.update(a={"must_understand": False}) + assert updated.extra_fields == { + "a": {"must_understand": False}, + "b": {"must_understand": False}, + } # --- V2 model -------------------------------------------------------------- @@ -495,8 +684,17 @@ def test_partial_keys_match_settable_model_fields() -> None: def test_v2_partial_keys_match_settable_model_fields() -> None: """The v2 partial TypedDict must list exactly the settable fields.""" - settable = {f.name for f in dataclasses.fields(ZarrV2ArrayMetadata) if f.init} - assert set(ZarrV2ArrayMetadataPartial.__annotations__) == settable + assert set(ZarrV2ArrayMetadataUpdate.__annotations__) == { + "shape", + "dtype", + "chunks", + "fill_value", + "order", + "compressor", + "filters", + "dimension_separator", + "attributes", + } def test_v2_to_key_value_splits_zarray_and_zattrs() -> None: @@ -555,22 +753,17 @@ def test_arrays_to_tuples(case: Expect[object, object]) -> None: def test_v3_from_json_reconstructs_required_fields() -> None: """V3 from_json reconstructs the required fields from a document.""" doc = ZarrV3ArrayMetadata.create_default( - shape=(7,), - attributes={"a": 1}, - data_type=ZarrV3NamedConfig(name="int32", configuration={}), - codecs=(ZarrV3NamedConfig(name="bytes", configuration={"endian": "little"}),), + shape=(7,), attributes={"a": 1}, data_type="int32" ).to_json() model = ZarrV3ArrayMetadata.from_json(doc) assert model.shape == (7,) - assert model.data_type == ZarrV3NamedConfig(name="int32", configuration={}) + assert model.data_type.to_json() == "int32" assert model.attributes == {"a": 1} def test_v3_from_json_defaults_for_omitted_optionals() -> None: """V3 from_json supplies defaults for omitted optional fields.""" - doc = ZarrV3ArrayMetadata.create_default( - attributes={}, storage_transformers=(), dimension_names=UNSET - ).to_json() + doc = ZarrV3ArrayMetadata.create_default(attributes={}, storage_transformers=()).to_json() # to_json omits these entirely; from_json must restore defaults model = ZarrV3ArrayMetadata.from_json(doc) assert model.attributes == {} @@ -580,9 +773,7 @@ def test_v3_from_json_defaults_for_omitted_optionals() -> None: def test_v3_from_json_routes_unknown_keys_to_extra_fields() -> None: """V3 from_json routes unknown top-level keys into extra_fields.""" - doc = ZarrV3ArrayMetadata.create_default( - extra_fields={"my_ext": {"must_understand": False}} - ).to_json() + doc = ZarrV3ArrayMetadata.create_default(my_ext={"must_understand": False}).to_json() model = ZarrV3ArrayMetadata.from_json(doc) assert model.extra_fields == {"my_ext": {"must_understand": False}} @@ -648,19 +839,14 @@ def test_from_key_value_missing_key_raises( shape=(10,), attributes={"a": 1}, dimension_names=("x",), - storage_transformers=(ZarrV3NamedConfig(name="t", configuration={}),), - extra_fields={"ext": {"must_understand": False}}, + storage_transformers=({"name": "acme.t"},), + ext={"must_understand": False}, ), id="v3-full", ), pytest.param( ZarrV3ArrayMetadata, - ZarrV3ArrayMetadata.create_default( - attributes={}, - dimension_names=UNSET, - storage_transformers=(), - extra_fields={}, - ), + ZarrV3ArrayMetadata.create_default(attributes={}, storage_transformers=()), id="v3-empty-optionals", ), pytest.param( @@ -717,18 +903,21 @@ def test_v2_roundtrip_json_model_json() -> None: ZarrV3ArrayMetadata.create_default( shape=(2,), attributes={"a": {"b": [1]}}, - # A name nothing in the scope claims, so reading it back judges only + # A name nothing in the scope claims, so reading it judges only # the document, and its configuration can nest. - codecs=(ZarrV3NamedConfig(name="acme.nested", configuration={"opts": {"level": 1}}),), - extra_fields={"ext": {"must_understand": False, "cfg": {"x": [1]}}}, + codecs=({"name": "acme.nested", "configuration": {"opts": {"level": 1}}},), + ext={"must_understand": False, "cfg": {"x": [1]}}, ), id="v3", ), pytest.param( ZarrV2ArrayMetadata.create_default( attributes={"a": {"b": [1]}}, - compressor={"id": "zstd", "opts": {"level": 1}}, - filters=({"id": "delta", "cfg": [1]},), + # Ids nothing in the scope claims, so their parameters can nest; + # a complex type, whose fill value is a pair. + dtype=" None: doc["data_type"] = "int32" doc["codecs"] = ({"name": "bytes", "configuration": {"endian": "little"}},) model = ZarrV3ArrayMetadata.from_json(doc) - assert model.data_type == ZarrV3NamedConfig(name="int32", configuration={}) + assert (model.data_type.json, model.data_type.name) == ("int32", "int32") assert model.to_json()["data_type"] == "int32" -@pytest.mark.parametrize("name", ["bytes", "ANY string", "urn:example:codec"]) -def test_metadata_field_accepts_any_string_name(name: str) -> None: - """The structural layer checks the name type, not syntax or registration.""" +@pytest.mark.parametrize("name", ["bytes", "acme.codec", "urn:example:codec"]) +def test_metadata_field_accepts_a_name_as_the_spec_names_one(name: str) -> None: + """The structural layer checks the name is one the spec gives an extension, not that anything registered it.""" assert validate_metadata_field_v3({"name": name}) == () @pytest.mark.parametrize("value", [0, 1, "false", None]) def test_metadata_field_must_understand_must_be_boolean(value: object) -> None: """must_understand is a JSON boolean, not a truthy scalar.""" - problems = validate_metadata_field_v3({"name": "x", "must_understand": value}) + problems = validate_metadata_field_v3({"name": "acme.x", "must_understand": value}) assert [(problem.loc, problem.kind) for problem in problems] == [ (("must_understand",), "invalid_type") ] @@ -787,8 +976,8 @@ def test_metadata_field_must_understand_must_be_boolean(value: object) -> None: def test_metadata_field_rejects_unknown_envelope_member() -> None: """Unknown envelope keys cannot be silently discarded during normalization.""" - problems = validate_metadata_field_v3({"name": "x", "typo": 1}) - assert [(problem.loc, problem.kind) for problem in problems] == [(("typo",), "invalid_value")] + problems = validate_metadata_field_v3({"name": "acme.x", "typo": 1}) + assert [(problem.loc, problem.kind) for problem in problems] == [(("typo",), "unknown_key")] @pytest.mark.parametrize("field", ["codecs", "storage_transformers"]) @@ -830,11 +1019,12 @@ def test_v2_roundtrip_with_compressor_and_filters() -> None: # Non-None compressor/filters must round-trip; extra assertion on .compressor. """A v2 model with non-None compressor and filters round-trips.""" compressor: ZarrV2CodecMetadata = {"id": "blosc", "clevel": 5} - filters: tuple[ZarrV2CodecMetadata, ...] = ({"id": "delta"},) + filters: tuple[ZarrV2CodecMetadata, ...] = ({"id": "delta", "dtype": " None: doc = ZarrV2ArrayMetadata.create_default(shape=(4,), attributes={"a": 1}, dtype=" None: ] -def test_v2_from_key_value_ignores_zarray_extra_members() -> None: - """Other raw `.zarray` members "SHOULD be ignored by implementations" (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L91-L92).""" +def test_v2_from_key_value_keeps_zarray_extra_members() -> None: + """Other raw `.zarray` members "SHOULD be ignored by implementations" (https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L91-L92): left unjudged, and kept as written.""" doc: dict[str, object] = dict(ZarrV2ArrayMetadata.create_default().to_json()) doc.pop("attributes", None) doc["vendor_extension"] = {} model = ZarrV2ArrayMetadata.from_key_value({".zarray": json.dumps(doc).encode()}) - assert "vendor_extension" not in model.to_json() + assert model.to_json()["vendor_extension"] == {} + assert model.extra_fields == {"vendor_extension": {}} def test_v2_zattrs_presence_round_trips() -> None: @@ -1038,7 +1229,7 @@ def test_validate_json_reports_json_in_message() -> None: METADATA_FIELD_VALIDATE_CASES: list[Expect[object, frozenset[tuple[str | int, ...]]]] = [ Expect("bytes", frozenset(), id="bare-string"), - Expect({"name": "x", "configuration": {"a": 1}}, frozenset(), id="named-config"), + Expect({"name": "acme.x", "configuration": {"a": 1}}, frozenset(), id="named-config"), Expect({"name": "bytes"}, frozenset(), id="name-only"), Expect(5, frozenset({()}), id="not-str-or-mapping"), Expect({"configuration": {}}, frozenset({("name",)}), id="missing-name"), @@ -1091,11 +1282,11 @@ def test_parse_metadata_field_v3( # invalid cases (a subset check, so accumulation of OTHER problems is allowed). -def _build_v3(**overrides: Unpack[ZarrV3ArrayMetadataPartial]) -> dict[str, object]: +def _build_v3(**overrides: Unpack[ZarrV3ArrayMetadataJSONPartial]) -> dict[str, object]: return dict(ZarrV3ArrayMetadata.create_default(**overrides).to_json()) -def _build_v2(**overrides: Unpack[ZarrV2ArrayMetadataPartial]) -> dict[str, object]: +def _build_v2(**overrides: Unpack[ZarrV2ArrayMetadataUpdate]) -> dict[str, object]: return dict(ZarrV2ArrayMetadata.create_default(**overrides).to_json()) @@ -1124,7 +1315,7 @@ def _set(key: str, value: object) -> Callable[[dict], object]: id="valid-with-attributes-and-dim-names", ), Expect( - lambda: _build_v3(extra_fields={"my_ext": {"must_understand": False}}), + lambda: _build_v3(my_ext={"must_understand": False}), frozenset(), id="valid-with-extra-fields", ), @@ -1247,17 +1438,12 @@ def test_array_metadata_guards( ExpectFail(lambda: {"zarr_format": 2}, MetadataValidationError, id="x"), id="v2-missing-required", ), - pytest.param( - ZarrV3NamedConfig, - ExpectFail(lambda: 5, MetadataValidationError, id="x"), - id="zarr-metadata-bad-input", - ), ] @pytest.mark.parametrize(("model", "case"), FROM_JSON_REJECT_PARAMS) def test_from_json_rejects_malformed( - model: type[ZarrV3ArrayMetadata | ZarrV2ArrayMetadata | ZarrV3NamedConfig], + model: type[ZarrV3ArrayMetadata | ZarrV2ArrayMetadata], case: ExpectFail[Callable[[], object]], ) -> None: """from_json raises MetadataValidationError on a malformed document.""" @@ -1408,7 +1594,7 @@ def test_v2_dtype_must_be_string_or_records() -> None: ( validate_array_metadata_v2, {**ZarrV2ArrayMetadata.create_default().to_json(), "dtype": (("f0", b""),)}, - ("dtype",), + ("dtype", 0, 1), ), ( validate_array_metadata_v2, @@ -1431,15 +1617,18 @@ def test_error_bytes_are_not_an_array( def test_v2_structured_dtype_records_accepted() -> None: """A structured v2 dtype (field records, optionally nested/shaped) validates.""" dtype = (("a", " None: - """A field record with the wrong arity is rejected.""" + """A field record with the wrong arity is rejected, at the record.""" doc = dict(ZarrV2ArrayMetadata.create_default().to_json()) | {"dtype": (("a",),)} problems = validate_array_metadata_v2(doc) - assert [p.loc for p in problems] == [("dtype",)] + assert [p.loc for p in problems] == [("dtype", "fields", 0)] def test_v2_order_literal_enforced() -> None: @@ -1457,18 +1646,18 @@ def test_v2_compressor_must_be_codec_or_none() -> None: def test_v2_compressor_requires_string_id() -> None: - """A compressor mapping without a string id is rejected.""" + """A compressor mapping without a string id is rejected, at the id.""" doc = dict(ZarrV2ArrayMetadata.create_default().to_json()) | {"compressor": {"level": 3}} problems = validate_array_metadata_v2(doc) - assert [p.loc for p in problems] == [("compressor",)] + assert [p.loc for p in problems] == [("compressor", "id")] def test_v2_filters_must_be_codec_sequence_or_none() -> None: - """Filters that are not null or a sequence of codec configs are rejected.""" - for bad in (7, (5,), "gzip"): + """Filters that are not null or a sequence of codec configs are rejected: an item that is no codec at the item, anything else at the field.""" + for bad, at in ((7, ("filters",)), ((5,), ("filters", 0)), ("gzip", ("filters",))): doc = dict(ZarrV2ArrayMetadata.create_default().to_json()) | {"filters": bad} problems = validate_array_metadata_v2(doc) - assert [(p.loc, p.kind) for p in problems] == [(("filters",), "invalid_type")], bad + assert [(p.loc, p.kind) for p in problems] == [(at, "invalid_type")], bad def test_v2_shape_and_chunks_must_have_equal_rank() -> None: @@ -1538,15 +1727,17 @@ def test_array_zarr_format_rejects_float( assert [(p.loc, p.kind) for p in validate(document)] == [(("zarr_format",), "invalid_value")] -def test_array_v2_ignores_unknown_document_member() -> None: - """Other .zarray keys "SHOULD NOT be present ... and SHOULD be ignored": tolerated, dropped. +def test_array_v2_keeps_an_unknown_document_member() -> None: + """Other .zarray keys "SHOULD NOT be present ... and SHOULD be ignored": tolerated, left unjudged, and kept as written, in `extra_fields`. https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L91-L92 """ doc = dict(ZarrV2ArrayMetadata.create_default().to_json()) | {"unexpected": 1} assert validate_array_metadata_v2(doc) == () - assert "unexpected" not in ZarrV2ArrayMetadata.from_json(doc).to_json() + model = ZarrV2ArrayMetadata.from_json(doc) + assert model.to_json()["unexpected"] == 1 + assert model.extra_fields == {"unexpected": 1} @pytest.mark.parametrize( @@ -1596,12 +1787,218 @@ def test_from_key_value_rejects_non_standard_json_constant() -> None: ZarrV3ArrayMetadata.from_key_value({"zarr.json": raw.encode()}) -def test_to_key_value_rejects_non_finite_model_value() -> None: - """Strict encoding prevents directly-constructed models from writing invalid JSON.""" - model = ZarrV3ArrayMetadata.create_default(fill_value=float("nan")) +@pytest.mark.parametrize( + ("members", "problems"), + [ + ({"fill_value": math.nan}, [(("fill_value",), "invalid_value")]), + # The default fill value, 0, is not a boolean: a data type goes with + # a fill value of it. + ({"data_type": "bool"}, [(("fill_value",), "invalid_type")]), + # Values of several bytes take an `endian`. + ( + {"data_type": "int16", "codecs": ({"name": "bytes"},)}, + [(("codecs", 0, "configuration", "endian"), "missing_key")], + ), + ({"shape": (2,), "dimension_names": ("x", "y")}, [(("dimension_names",), "invalid_value")]), + # The shape is not derived from a grid, which the scalar default + # shape does not fit: a grid goes with the shape it fits. + ( + {"chunk_grid": {"name": "regular", "configuration": {"chunk_shape": (10, 10)}}}, + [(("chunk_grid", "configuration", "chunk_shape"), "invalid_value")], + ), + ], + ids=[ + "non-finite-fill-value", + "fill-value-of-another-type", + "no-endian", + "names-past-shape", + "grid-of-another-rank", + ], +) +def test_error_create_default_refuses_a_document_with_a_problem( + members: ZarrV3ArrayMetadataJSONPartial, problems: list[tuple[tuple[str | int, ...], str]] +) -> None: + """A model comes from a read, so a document its read refuses makes none, and nothing is written.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV3ArrayMetadata.create_default(**members) + assert [(p.loc, p.kind) for p in raised.value.problems] == problems - with pytest.raises(MetadataValidationError, match="fill_value: non-finite float nan"): - model.to_key_value() + +@pytest.mark.parametrize( + ("members", "problems"), + [ + ({"fill_value": 300}, [(("fill_value",), "invalid_value")]), + # A shape goes with a grid that fits it. + ({"shape": (4, 4)}, [(("chunk_grid", "configuration", "chunk_shape"), "invalid_value")]), + ], + ids=["fill-value-out-of-range", "shape-without-its-grid"], +) +def test_error_update_refuses_a_document_with_a_problem( + members: ZarrV3ArrayMetadataJSONPartial, problems: list[tuple[tuple[str | int, ...], str]] +) -> None: + base = ZarrV3ArrayMetadata.create_default(shape=(4,)) + with pytest.raises(MetadataValidationError) as raised: + base.update(**members) + assert [(p.loc, p.kind) for p in raised.value.problems] == problems + + +def test_error_update_refuses_to_leave_out_a_member_a_document_holds() -> None: + base = ZarrV3ArrayMetadata.create_default(shape=(4,)) + with pytest.raises(MetadataValidationError) as raised: + base.update(shape=UNSET) # pyright: ignore[reportArgumentType] + assert [(p.loc, p.kind) for p in raised.value.problems] == [(("shape",), "missing_key")] + + +@pytest.mark.parametrize( + ("shape", "kind"), + [ + (5, "invalid_type"), + (None, "invalid_type"), + (("a",), "invalid_type"), + ((-1,), "invalid_value"), + ], + ids=["a-number", "null", "not-integers", "negative"], +) +def test_error_create_default_reports_a_shape_it_cannot_read(shape: object, kind: str) -> None: + """As the read reports it, rather than failing to derive a grid from it.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV3ArrayMetadata.create_default(shape=shape) # pyright: ignore[reportArgumentType] + assert [(p.loc, p.kind) for p in raised.value.problems] == [(("shape",), kind)] + + +def test_a_codec_that_is_not_json_is_placed_by_the_definition_that_claims_its_name() -> None: + """Its kind is that definition's, as a codec refused for a configuration that is JSON takes it: `gzip` is a bytes -> bytes codec, so the pipeline still lacks an array -> bytes one.""" + document = { + **ZarrV3ArrayMetadata.create_default().to_json(), + "codecs": [{"name": "gzip", "configuration": {"level": math.nan}}], + } + assert [(p.loc, p.kind) for p in validate_array_metadata_v3(document)] == [ + (("codecs", 0, "configuration", "level"), "invalid_value"), + (("codecs",), "invalid_value"), + ] + + +def test_a_shard_nested_as_deep_as_a_reader_walks_is_read_and_written() -> None: + """Every walker of fields takes more than one frame per shard, so the deepest nesting the cap admits is where the interpreter's limit would show; one shard deeper is the depth problem.""" + little = {"name": "bytes", "configuration": {"endian": "little"}} + + def nested(shards: int) -> list[object]: + codecs: list[object] = [little] + for _ in range(shards): + codecs = [ + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [1], + "codecs": codecs, + "index_codecs": [little], + }, + } + ] + return codecs + + # Each shard is three levels -- its object, its configuration and the + # `codecs` in it -- and the innermost codec's configuration is the + # last container a reader walks. + deepest = (JSON_DEPTH - 4) // 3 + codecs = nested(deepest) + document = {**ZarrV3ArrayMetadata.create_default(shape=(2,)).to_json(), "codecs": codecs} + assert validate_array_metadata_v3(document) == () + model = ZarrV3ArrayMetadata.from_json(document) + assert json_text(model.to_json()) == json_text(cast("JSONValue", document)) + assert ZarrV3ArrayMetadata.from_key_value(model.to_key_value()) == model + assert hash(model) == hash(ZarrV3ArrayMetadata.from_json(document)) + assert pickle.loads(pickle.dumps(model)) == model + assert copy.deepcopy(model) == model + (shard,) = model.codecs + assert json_text(canonical_of(shard, ())) == json_text(cast("JSONValue", codecs[0])) + # A shard and its index codec at each level, and the innermost codec. + assert len(list(with_problems(fields_of(shard), ()))) == 2 * deepest + 1 + problems = validate_array_metadata_v3({**document, "codecs": nested(deepest + 1)}) + assert {(problem.kind, len(problem.loc), problem.message) for problem in problems} == { + ("invalid_value", JSON_DEPTH, f"nested deeper than the {JSON_DEPTH} levels a reader walks") + } + + +def test_a_model_is_written_as_deep_as_it_is_read() -> None: + """A fill value as deep as a reader walks, of a data type nothing in scope claims, which takes any JSON; pickled and deep-copied too, which take two frames a level.""" + fill_value: dict[str, object] = {} + for _ in range(JSON_DEPTH - 2): + fill_value = {"x": fill_value} + document = { + **ZarrV3ArrayMetadata.create_default().to_json(), + "data_type": "acme.deep", + "fill_value": fill_value, + } + model = ZarrV3ArrayMetadata.from_json(document) + assert model.to_json()["fill_value"] == fill_value + assert ZarrV3ArrayMetadata.from_key_value(model.to_key_value()) == model + assert pickle.loads(pickle.dumps(model)) == model + assert copy.deepcopy(model) == model + + +def test_a_field_s_configuration_member_is_counted_from_the_field_s_root() -> None: + # `validate_metadata_field_v3` judges a field alone, at its own root: a + # member of its configuration sits two levels down, and the cap counts + # from the root, not from the member. + nested: dict[str, object] = {} + for _ in range(JSON_DEPTH - 2): + nested = {"x": nested} + problems = validate_metadata_field_v3({"name": "acme.x", "configuration": {"y": nested}}) + assert [(len(p.loc), p.kind) for p in problems] == [(JSON_DEPTH, "invalid_value")] + assert problems[0].loc[:2] == ("configuration", "y") + shallower = {"name": "acme.x", "configuration": {"y": nested["x"]}} + assert validate_metadata_field_v3(shallower) == () + + +def test_a_v2_dtype_of_nested_records_is_read_to_the_levels_a_reader_walks() -> None: + """Field records nest a dtype two levels a record: the shape check recursed a record at a time with no cap, and a thousand records overflowed.""" + + def records(levels: int) -> object: + dtype: object = " None: + """The v2 models copied with `copy.deepcopy`, two frames a level, and overflowed on documents their validators accept.""" + fill_value: dict[str, object] = {} + for _ in range(JSON_DEPTH - 2): + fill_value = {"x": fill_value} + # An object type, whose fill value is any JSON. + document = { + **ZarrV2ArrayMetadata.create_default(shape=(2,), dtype="|O").to_json(), + "fill_value": fill_value, + } + assert validate_array_metadata_v2(document) == () + model = ZarrV2ArrayMetadata.from_json(document) + assert json_text(model.to_json()) == json_text(cast("JSONValue", document)) + assert ZarrV2ArrayMetadata.from_key_value(model.to_key_value()) == model + assert hash(model) == hash(ZarrV2ArrayMetadata.from_json(document)) + assert pickle.loads(pickle.dumps(model)) == model + assert copy.deepcopy(model) == model + problems = validate_array_metadata_v2({**document, "fill_value": {"x": fill_value}}) + assert [(problem.kind, len(problem.loc)) for problem in problems] == [ + ("invalid_value", JSON_DEPTH) + ] def test_v3_node_type_literal_enforced() -> None: @@ -1656,25 +2053,6 @@ def test_from_key_value_missing_key_kind() -> None: ) -def test_extra_fields_overlap_raises_metadata_error() -> None: - """The extra-fields overlap invariant raises MetadataValidationError (a ValueError).""" - with pytest.raises(MetadataValidationError, match="Extra fields") as exc_info: - ZarrV3ArrayMetadata.create_default(extra_fields={"shape": {"must_understand": False}}) - assert [p.kind for p in exc_info.value.problems] == ["invalid_value"] - - -def test_extension_point_fields_annotated_with_role_alias() -> None: - """Extension-point fields are annotated with ZarrV3MetadataField (the - logical role), not ZarrV3NamedConfig (the current serialized form), so a - future widening of the field union does not move annotation sites.""" - assert ZarrV3MetadataField is ZarrV3NamedConfig - annotations = ZarrV3ArrayMetadata.__annotations__ - for field_name in ("data_type", "chunk_grid", "chunk_key_encoding"): - assert annotations[field_name] == "ZarrV3MetadataField" - for field_name in ("codecs", "storage_transformers"): - assert annotations[field_name] == "tuple[ZarrV3MetadataField, ...]" - - # --- Adversarial-probe fixes: documents that used to pass validation --------- @@ -1806,7 +2184,8 @@ def test_array_guards_reject_noncanonical_nested_json() -> None: # Raw bits of 16, whose fill value is two byte values. v3 = dict(ZarrV3ArrayMetadata.create_default().to_json()) | {"data_type": "r16"} v3["fill_value"] = range(2) - v2 = dict(ZarrV2ArrayMetadata.create_default().to_json()) + # A complex type, whose fill value is a pair. + v2 = dict(ZarrV2ArrayMetadata.create_default(dtype=" None: https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L1575-L1578 """ model = ZarrV3ArrayMetadata.create_default( - extra_fields={ - "ext_a": {"name": "a", "must_understand": False}, - "ext_b": {"name": "b"}, - "ext_c": {"name": "c", "must_understand": True}, - "ext_d": 123, - } + ext_a={"name": "a", "must_understand": False}, + ext_b={"name": "b"}, + ext_c={"name": "c", "must_understand": True}, + ext_d=123, ) assert set(model.must_understand_fields) == {"ext_b", "ext_c", "ext_d"} recognized = {"ext_b"} @@ -1841,9 +2218,7 @@ def test_must_understand_fields_partition() -> None: def test_must_understand_fields_empty_when_all_waived() -> None: """must_understand_fields is empty when every extra field is explicitly waived.""" - model = ZarrV3ArrayMetadata.create_default( - extra_fields={"ext_a": {"name": "a", "must_understand": False}} - ) + model = ZarrV3ArrayMetadata.create_default(ext_a={"name": "a", "must_understand": False}) assert model.must_understand_fields == {} @@ -1858,10 +2233,10 @@ def test_dimension_names_null_field_rejected() -> None: doc = dict(ZarrV3ArrayMetadata.create_default().to_json()) | {"dimension_names": None} problems = validate_array_metadata_v3(doc) assert [(p.loc, p.kind) for p in problems] == [(("dimension_names",), "invalid_type")] - # and the model's own None spelling correctly maps to key absence - assert ( - "dimension_names" not in ZarrV3ArrayMetadata.create_default(dimension_names=UNSET).to_json() - ) + # and the model's own UNSET spelling correctly maps to key absence + model = ZarrV3ArrayMetadata.create_default() + assert model.dimension_names is UNSET + assert "dimension_names" not in model.to_json() # --- create_default derives the chunk grid from shape ------------------------ @@ -1872,16 +2247,17 @@ def test_v3_create_default_chunk_grid_follows_shape() -> None: one chunk covering the array (chunk_shape == shape), instead of silently keeping the scalar default's 0-d grid.""" model = ZarrV3ArrayMetadata.create_default(shape=(100, 100)) - assert model.chunk_grid == ZarrV3NamedConfig( - name="regular", configuration={"chunk_shape": (100, 100)} - ) + assert model.chunk_grid.to_json() == { + "name": "regular", + "configuration": {"chunk_shape": (100, 100)}, + } def test_v3_create_default_explicit_chunk_grid_respected() -> None: """An explicit chunk_grid override wins over the shape-derived default.""" - grid = ZarrV3NamedConfig(name="regular", configuration={"chunk_shape": (10, 10)}) + grid: ZarrV3NamedConfigJSON = {"name": "regular", "configuration": {"chunk_shape": (10, 10)}} model = ZarrV3ArrayMetadata.create_default(shape=(100, 100), chunk_grid=grid) - assert model.chunk_grid == grid + assert model.chunk_grid.to_json() == grid def test_v2_create_default_chunks_follow_shape() -> None: @@ -1901,24 +2277,13 @@ def test_v3_create_default_zero_length_dimensions() -> None: the regular grid asks for chunk sizes greater than zero, so the written grid is one every reader takes, and it fits the shape it chunks.""" model = ZarrV3ArrayMetadata.create_default(shape=(0, 3)) - assert model.chunk_grid.configuration["chunk_shape"] == (1, 3) + assert model.chunk_grid.to_json() == { + "name": "regular", + "configuration": {"chunk_shape": (1, 3)}, + } assert validate_array_metadata_v3(model.to_json()) == () -def test_create_default_derivation_is_one_way() -> None: - """Overriding the chunk grid (v3) or chunks (v2) without shape leaves the - scalar default shape=() untouched: a user-supplied chunk_grid is an - extension point taken verbatim, and deriving shape from it would require - interpreting grid configurations, which the model layer never does.""" - grid = ZarrV3NamedConfig(name="regular", configuration={"chunk_shape": (10, 10)}) - v3 = ZarrV3ArrayMetadata.create_default(chunk_grid=grid) - assert v3.shape == () - assert v3.chunk_grid == grid - v2 = ZarrV2ArrayMetadata.create_default(chunks=(10, 10)) - assert v2.shape == () - assert v2.chunks == (10, 10) - - # --- v2 dimension_separator default (roborev job 426) ------------------------- @@ -1953,7 +2318,7 @@ def test_v2_null_dimension_separator_rejected() -> None: document grammar has no null spelling for this field.""" doc = dict(ZarrV2ArrayMetadata.create_default().to_json()) | {"dimension_separator": None} problems = validate_array_metadata_v2(doc) - assert [(p.loc, p.kind) for p in problems] == [(("dimension_separator",), "invalid_value")] + assert [(p.loc, p.kind) for p in problems] == [(("dimension_separator",), "invalid_type")] def test_dimension_names_absent_and_all_null_are_distinct() -> None: @@ -1974,3 +2339,133 @@ def test_dimension_names_absent_and_all_null_are_distinct() -> None: assert "dimension_names" not in absent.to_json() assert absent.to_json() == absent_doc assert explicit.to_json() == explicit_doc + + +@pytest.mark.parametrize( + ("validate", "document", "member", "value", "kind"), + [ + (validate_array_metadata_v3, ZarrV3ArrayMetadata, "zarr_format", "3", "invalid_type"), + (validate_array_metadata_v3, ZarrV3ArrayMetadata, "zarr_format", 2, "invalid_value"), + (validate_array_metadata_v3, ZarrV3ArrayMetadata, "node_type", 5, "invalid_type"), + (validate_array_metadata_v3, ZarrV3ArrayMetadata, "node_type", "group", "invalid_value"), + (validate_array_metadata_v2, ZarrV2ArrayMetadata, "order", 1, "invalid_type"), + (validate_array_metadata_v2, ZarrV2ArrayMetadata, "order", "Q", "invalid_value"), + ( + validate_array_metadata_v2, + ZarrV2ArrayMetadata, + "dimension_separator", + ":", + "invalid_value", + ), + ], +) +def test_error_a_member_outside_the_values_it_takes( + validate: Callable[[object], tuple[ValidationProblem, ...]], + document: type[ZarrV3ArrayMetadata | ZarrV2ArrayMetadata], + member: str, + value: object, + kind: str, +) -> None: + # Of the wrong type when none of the values it takes is of its JSON + # type -- a string where a number belongs -- else of the wrong value. + written = {**document.create_default().to_json(), member: value} + assert [(p.loc, p.kind) for p in validate(written)] == [((member,), kind)] + + +def test_error_a_document_nested_deeper_than_a_reader_walks_is_a_problem() -> None: + # Not a `RecursionError`: a 10 KB document a stranger wrote, nested past + # the interpreter's limit, is refused at the level past the last a + # reader walks, by every validator, guard and reader. + deep: dict[str, object] = {} + for _ in range(2_000): + deep = {"a": deep} + document = {**ZarrV3ArrayMetadata.create_default().to_json(), "attributes": deep} + problems = validate_array_metadata_v3(document) + assert [(len(p.loc), p.kind) for p in problems] == [(JSON_DEPTH, "invalid_value")] + assert problems[0].loc[:2] == ("attributes", "a") + assert not is_json(document) + assert not is_array_metadata_v3(document) + assert not is_group_metadata_v3(document) + assert not is_array_metadata_v2(document) + assert not is_group_metadata_v2(document) + with pytest.raises(MetadataValidationError): + ZarrV3ArrayMetadata.from_json(document) + + +def test_error_a_field_object_in_a_document_is_not_json() -> None: + # A `AcceptedField` built by hand, with a configuration its definition refuses, + # smuggled into a document: refused as what it is, so nothing built by + # hand passes as read. A model holds its own fields as read. + smuggled = AcceptedField( + json="gzip", name="gzip", definition=GZIP_CODEC, configuration={"level": 99, "window": 1} + ) + document = { + **ZarrV3ArrayMetadata.create_default(shape=(2,)).to_json(), + "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}, smuggled], + } + assert [(p.loc, p.kind) for p in validate_array_metadata_v3(document)] == [ + (("codecs", 1), "invalid_type") + ] + assert not is_array_metadata_v3(document) + with pytest.raises(MetadataValidationError): + ZarrV3ArrayMetadata.from_json(document) + # Nor does `update` take one among the members it is given: a model's + # own fields are no exception, since `update` reads JSON. + model = ZarrV3ArrayMetadata.create_default(shape=(2,)) + with pytest.raises(MetadataValidationError) as raised: + model.update(codecs=cast("Any", (model.codecs[0].to_json(), smuggled))) + assert [(p.loc, p.kind) for p in raised.value.problems] == [(("codecs", 1), "invalid_type")] + with pytest.raises(MetadataValidationError): + model.update(codecs=(cast("Any", model.codecs[0]),)) + + +def test_a_problem_shows_an_integer_too_long_to_write_by_its_size( + interpreter_writes_4300_digits: None, +) -> None: + # `int` refuses to write more than 4,300 digits; a message says the + # size instead of raising. + document = {**ZarrV3ArrayMetadata.create_default().to_json(), "zarr_format": 10**5000} + (problem,) = validate_array_metadata_v3(document) + assert problem.message == f"expected 3, got an integer of {(10**5000).bit_length()} bits" + (problem,) = validate_metadata_field_v3({"name": "gzip", 10**5000: 1}) + assert problem.message == ( + f"non-string metadata field key an integer of {(10**5000).bit_length()} bits" + ) + # Held by a container, JSON or not, or by a key, it is what the + # interpreter will not write. + for held in ([10**5000], [10**5000, object()]): + document = {**ZarrV3ArrayMetadata.create_default().to_json(), "zarr_format": held} + (problem,) = validate_array_metadata_v3(document) + assert ( + problem.message == "expected 3, got a value of type list the interpreter will not write" + ) + (problem,) = validate_metadata_field_v3({"name": "gzip", (10**5000,): 1}) + assert problem.message == ( + "non-string metadata field key a value of type tuple the interpreter will not write" + ) + # So is what `refine_json` reports as no JSON at all: a set holding one. + document = {**ZarrV3ArrayMetadata.create_default().to_json(), "attributes": {"a": {10**5000}}} + (problem,) = validate_array_metadata_v3(document) + assert (problem.loc, problem.message) == ( + ("attributes", "a"), + "not a JSON-serializable value: a value of type set the interpreter will not write", + ) + + +def test_a_problem_shows_a_value_nested_too_deep_to_write_by_saying_so() -> None: + # `_refine` stops at `JSON_DEPTH`, and a value nested past it is said to + # be too deep, not shown by a repr, whose own limit the interpreter and + # platform set: one level past, which any repr would write, is enough. + deep: list[object] = [] + innermost = deep + for _ in range(JSON_DEPTH + 1): + nested: list[object] = [] + innermost.append(nested) + innermost = nested + document = {**ZarrV3ArrayMetadata.create_default().to_json(), "zarr_format": deep} + # One problem: a value the literal check refuses is not walked, so the + # depth rule does not judge it too. + assert [ + (problem.loc, problem.kind, problem.message) + for problem in validate_array_metadata_v3(document) + ] == [(("zarr_format",), "invalid_type", "expected 3, got a value nested too deep to show")] diff --git a/packages/zarr-metadata/tests/model/test_consolidating.py b/packages/zarr-metadata/tests/model/test_consolidating.py new file mode 100644 index 0000000000..eebbd3ad7c --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_consolidating.py @@ -0,0 +1,341 @@ +"""A group's consolidated metadata built from node models: each accepted when it reads the same in the group's scope, or gains there; refused when it conflicts or would lose.""" + +from __future__ import annotations + +import dataclasses +import pickle +from typing import Any, cast, get_type_hints + +import pytest + +from zarr_metadata._json import JSON_DEPTH +from zarr_metadata.model import ( + MetadataValidationError, + ZarrV3ArrayMetadata, + ZarrV3ConsolidatedMetadata, + ZarrV3ConsolidatedMetadataInput, + ZarrV3GroupMetadata, + ZarrV3GroupMetadataReading, + ZarrV3NodeMetadataInput, + is_group_metadata_v3, + parse_group_metadata_v3, + read_group_metadata_v3, + validate_group_metadata_v3, +) +from zarr_metadata.v3._scope import ( + claims_of, +) +from zarr_metadata.v3.codec.crc32c import Empty +from zarr_metadata.v3.codec.gzip import GZIP_CODEC +from zarr_metadata.v3.data_type.raw import RAW_BYTES_DATA_TYPE +from zarr_metadata.v3.definition import ( + CORE, + CORE_AND_EXTENSIONS, + CodecDefinition, + Context, +) + +ARRAY: dict[str, Any] = { + **ZarrV3ArrayMetadata.create_default(shape=(4,)).to_json(), + "codecs": ("bytes",), +} +"""A `uint8` array whose `bytes` takes no configuration, so a private `bytes` of none reads it too.""" +ZSTD = {"name": "zstd", "configuration": {"level": 3, "checksum": False}} +LOOSE_ZSTD = {"name": "zstd", "configuration": {"level": 30}} +WITH_ZSTD = {**ARRAY, "codecs": (*ARRAY["codecs"], ZSTD)} +WITH_LOOSE_ZSTD = {**ARRAY, "codecs": (*ARRAY["codecs"], LOOSE_ZSTD)} +MY_BYTES = CodecDefinition(name="bytes", configuration=Empty, kind="array_bytes", size="static") +PRIVATE = CORE.extended_with(MY_BYTES) +INLINE: dict[str, Any] = {"kind": "inline", "must_understand": False} + + +def _group(scope: Context | None = None, **entries: object) -> ZarrV3GroupMetadata: + member: Any = {**INLINE, "metadata": entries} + return ZarrV3GroupMetadata.create_default(context=scope, consolidated_metadata=member) + + +@pytest.mark.parametrize( + ("scope", "child", "beside", "gained"), + [ + (CORE_AND_EXTENSIONS, ZarrV3ArrayMetadata(ARRAY), {}, False), + (CORE_AND_EXTENSIONS, ZarrV3ArrayMetadata(ARRAY, context=CORE), {}, False), + (CORE_AND_EXTENSIONS, ZarrV3ArrayMetadata(WITH_ZSTD, context=CORE), {}, True), + ( + CORE, + ZarrV3GroupMetadata( + _group(CORE, x=ZarrV3ArrayMetadata(ARRAY, context=CORE)).to_json(), + context=Context.of(), + ), + {"a/x": ZarrV3ArrayMetadata(ARRAY, context=Context.of())}, + True, + ), + ( + CORE, + { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": {**INLINE, "metadata": {"x": ZarrV3ArrayMetadata(ARRAY)}}, + }, + {"a/x": ARRAY}, + False, + ), + ], + ids=["same-scope", "unused-definitions", "gain", "nested-models", "model-in-a-document"], +) +def test_a_model_is_accepted_as_a_consolidated_entry_when_it_refines_into_the_scope( + scope: Context, + child: ZarrV3ArrayMetadata | ZarrV3GroupMetadata | dict[str, Any], + beside: dict[str, Any], + gained: bool, +) -> None: + """A node model given where consolidated metadata lists a document -- at the top, or inside a document listed there -- is accepted when the group's scope reads every claim of it identically -- its reading is kept, not read again -- or claims what the model's scope left unclaimed, when its document is read again there; the group then holds it in its own scope, equal to the group built of the documents, and writes plain JSON. A listed group's own listing is listed flat beside it, as the convention asks.""" + group = _group(scope, a=child, **beside) + held = group.consolidated_metadata + assert isinstance(held, ZarrV3ConsolidatedMetadata) + node = held.metadata["a"] + assert node.context is group.context + if isinstance(child, dict): + child = ZarrV3GroupMetadata(_documents(child), context=scope) + assert cast("Any", node).refines(child) + assert (node == child) is not gained + if not gained and isinstance(child, ZarrV3ArrayMetadata): + assert isinstance(node, ZarrV3ArrayMetadata) + assert node.reading.pipeline is child.reading.pipeline + written = {path: _documents(entry) for path, entry in beside.items()} + assert group == _group(scope, a=child.to_json(), **written) + assert is_group_metadata_v3(group.to_json()) + + +def _documents(value: object) -> object: + """`value` with every node model in it replaced by its document, however deep.""" + if isinstance(value, (ZarrV3ArrayMetadata, ZarrV3GroupMetadata)): + return value.to_json() + if isinstance(value, dict): + return {key: _documents(item) for key, item in cast("dict[str, object]", value).items()} + return value + + +def test_documents_as_written_and_pickle() -> None: + """A group built of models writes each child's document as the child wrote it, and pickles as any group does.""" + child = ZarrV3ArrayMetadata({**ARRAY, "data_type": {"name": "uint8"}}) + group = _group(None, a=child) + written = cast("Any", group.to_json()["consolidated_metadata"]) + assert written["metadata"]["a"] == child.to_json() + assert pickle.loads(pickle.dumps(group)) == group + + +def test_error_a_model_that_conflicts_with_the_scope_is_refused_at_its_path() -> None: + """A model read with a private `bytes`, given to a group whose scope reads the core `bytes`, is a conflict: a problem at the entry's path, naming the kind and the name, and telling the two definitions apart, not a silent re-read.""" + child = ZarrV3ArrayMetadata(ARRAY, context=PRIVATE) + with pytest.raises(MetadataValidationError) as raised: + _group(CORE_AND_EXTENSIONS, a=child) + (problem,) = raised.value.problems + assert problem.loc == ("consolidated_metadata", "metadata", "a", "codecs", 0) + assert problem.kind == "invalid_value" + assert problem.message == ( + "expected a document read in the group's scope, got a model that reads the codec " + "'bytes' by CodecDefinition(name='bytes') of Empty, which the group's scope reads by " + "CodecDefinition(name='bytes') of BytesCodecConfiguration" + ) + + +def test_error_a_model_that_would_lose_a_meaning_is_refused_at_its_path() -> None: + """A model read where `zstd` is claimed, given to a group whose scope leaves it unclaimed, would lose what it reads: refused as a conflict is.""" + child = ZarrV3ArrayMetadata(WITH_ZSTD, context=CORE_AND_EXTENSIONS) + with pytest.raises(MetadataValidationError) as raised: + _group(CORE, a=child) + (problem,) = raised.value.problems + assert problem.loc == ("consolidated_metadata", "metadata", "a", "codecs", 1) + assert problem.message.endswith("which the group's scope leaves unclaimed") + + +def test_error_a_gain_that_surfaces_a_problem_is_reported_at_the_problem() -> None: + """A model whose scope left `zstd` unclaimed, given to a group whose scope claims it, is read again there: a level that definition refuses is a problem where it sits.""" + child = ZarrV3ArrayMetadata(WITH_LOOSE_ZSTD, context=CORE) + with pytest.raises(MetadataValidationError) as raised: + _group(CORE_AND_EXTENSIONS, a=child) + assert [problem.loc for problem in raised.value.problems] == [ + ("consolidated_metadata", "metadata", "a", "codecs", 1, "configuration", "level") + ] + + +def test_readers_see_models_as_the_constructor_does() -> None: + """`read_group_metadata_v3` and `validate_group_metadata_v3` given a document holding models read them as the constructor does: the reading holds the adopted model, and the validator reports a conflict.""" + fine = {**INLINE, "metadata": {"a": ZarrV3ArrayMetadata(ARRAY)}} + document = {"zarr_format": 3, "node_type": "group", "consolidated_metadata": fine} + reading = read_group_metadata_v3(document) + assert reading.problems == () + assert isinstance(reading.consolidated["a"].metadata, ZarrV3ArrayMetadata) + assert validate_group_metadata_v3(document) == () + + +def test_error_the_validator_reports_a_conflicting_model_as_the_constructor_does() -> None: + """`validate_group_metadata_v3` given a document holding a model the scope conflicts with reports the conflict at the entry's path, as the constructor refuses it.""" + clashing = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": { + **INLINE, + "metadata": {"a": ZarrV3ArrayMetadata(ARRAY, context=PRIVATE)}, + }, + } + assert [problem.loc for problem in validate_group_metadata_v3(clashing)] == [ + ("consolidated_metadata", "metadata", "a", "codecs", 0) + ] + + +def test_consolidated_metadata_given_whole_is_taken_as_its_models() -> None: + """A `ZarrV3ConsolidatedMetadata` given as the member is its models at their paths: `update(consolidated_metadata=other.consolidated_metadata)` carries them over.""" + source = _group(CORE_AND_EXTENSIONS, a=ZarrV3ArrayMetadata(ARRAY)) + target = ZarrV3GroupMetadata.create_default(attributes={"t": 1}).update( + consolidated_metadata=source.consolidated_metadata + ) + assert target.consolidated_metadata == source.consolidated_metadata + assert target.attributes == {"t": 1} + + +def test_children_of_different_scopes_consolidate_in_their_join() -> None: + """Children read in different scopes are consolidated in `Context.joined` of them: each refines into the join, and the group reads in it.""" + a = ZarrV3ArrayMetadata(ARRAY, context=CORE) + b = ZarrV3ArrayMetadata(WITH_ZSTD, context=CORE_AND_EXTENSIONS) + scope = Context.joined(a.context, b.context) + group = _group(scope, a=a, b=b) + assert group.context == CORE_AND_EXTENSIONS + held = group.consolidated_metadata + assert isinstance(held, ZarrV3ConsolidatedMetadata) + assert held.metadata["a"] == a + assert held.metadata["b"] == b + + +def test_update_keeps_a_models_scope_apart_from_the_groups() -> None: + """A model given to `update` keeps nothing of its own scope in the group: the group's `context` is the group's, and the child's is the group's too.""" + group = ZarrV3GroupMetadata.create_default(context=CORE) + updated = group.update( + consolidated_metadata={ + "kind": "inline", + "must_understand": False, + "metadata": {"a": ZarrV3ArrayMetadata(ARRAY)}, + } + ) + assert updated.context == CORE + held = updated.consolidated_metadata + assert isinstance(held, ZarrV3ConsolidatedMetadata) + assert held.metadata["a"].context == CORE + + +def test_parse_gives_json_for_a_document_holding_models() -> None: + """`parse_group_metadata_v3` of a document holding models gives JSON, each model as its document, which the guard then says yes to: the trio agree.""" + document = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": {**INLINE, "metadata": {"a": ZarrV3ArrayMetadata(ARRAY)}}, + } + parsed = parse_group_metadata_v3(document) + member = cast("Any", parsed["consolidated_metadata"]) + assert member["metadata"]["a"] == ZarrV3ArrayMetadata(ARRAY).to_json() + assert is_group_metadata_v3(parsed) + + +def test_error_a_refused_model_entrys_reading_is_of_the_groups_scope() -> None: + """A group model refused as an entry is read again in the group's scope for its reading, so the group's reading holds no field, and no model, of another scope: `claims_of` its fields finds one scope.""" + child = ZarrV3GroupMetadata( + _group(CORE_AND_EXTENSIONS, x=ZarrV3ArrayMetadata(WITH_ZSTD)).to_json(), + context=CORE_AND_EXTENSIONS, + ) + document = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": {**INLINE, "metadata": {"a": child, "a/x": WITH_ZSTD}}, + } + reading = read_group_metadata_v3(document, context=CORE) + nested = reading.consolidated["a"] + assert isinstance(nested, ZarrV3GroupMetadataReading) + assert nested.metadata is None + assert len(nested.problems) != 0 + assert nested.consolidated["x"].metadata is None + assert claims_of(reading.fields()) == claims_of( + read_group_metadata_v3(_documents(document), context=CORE).fields() + ) + for _, field in reading.fields(): + assert field.definition is None or field.definition in CORE.definitions() + + +def test_error_a_model_listed_too_deep_is_refused_where_it_sits() -> None: + """A model valid on its own, nested to the last level a reader walks, sits past the cap as a listed document: refused there, as its document would be, by the validator, the constructor and the reader alike.""" + deep: list[object] = [] + for _ in range(JSON_DEPTH - 3): + deep = [deep] + child = ZarrV3ArrayMetadata({**ARRAY, "attributes": {"x": deep}}) + document = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": {**INLINE, "metadata": {"a": child}}, + } + problems = validate_group_metadata_v3(document) + assert [problem.loc[:4] for problem in problems] == [ + ("consolidated_metadata", "metadata", "a", "attributes") + ] + with pytest.raises(MetadataValidationError): + ZarrV3GroupMetadata(document) + assert read_group_metadata_v3(document).metadata is None + + +def test_error_a_data_type_conflict_names_the_kind_and_the_written_name() -> None: + """A conflict over a data type names the kind in words and the name as the document writes it -- `r16`, filed under `r*` -- and says the two definitions differ when they print alike.""" + strict = dataclasses.replace( + RAW_BYTES_DATA_TYPE, fill_value_rules=lambda configuration, nested, value: iter(()) + ) + child = ZarrV3ArrayMetadata( + {**ARRAY, "data_type": "r16", "fill_value": [0, 0], "codecs": ("bytes",)}, + context=CORE.extended_with(strict), + ) + with pytest.raises(MetadataValidationError) as raised: + _group(CORE, a=child) + (problem,) = raised.value.problems + assert problem.loc == ("consolidated_metadata", "metadata", "a", "data_type") + assert problem.message == ( + "expected a document read in the group's scope, got a model that reads the data type " + "'r16' by another definition than the one the group's scope reads it by, " + "DataTypeDefinition(name='r*') of RawBytesConfiguration" + ) + + +def test_error_a_conflict_between_definitions_alike_says_they_differ() -> None: + """Two definitions of one name that read the same TypedDict, differing in their rules, are told apart in the message by saying so, since nothing else shows it.""" + strict = dataclasses.replace(GZIP_CODEC, rules=lambda configuration, nested: iter(())) + child = ZarrV3ArrayMetadata( + {**ARRAY, "codecs": ("bytes", {"name": "gzip", "configuration": {"level": 1}})}, + context=CORE.extended_with(strict), + ) + with pytest.raises(MetadataValidationError) as raised: + _group(CORE, a=child) + (problem,) = raised.value.problems + assert problem.message == ( + "expected a document read in the group's scope, got a model that reads the codec " + "'gzip' by another definition than the one the group's scope reads it by, " + "CodecDefinition(name='gzip') of GzipCodecConfiguration" + ) + + +def test_the_input_types_are_types_a_signature_can_hold() -> None: + """`ZarrV3NodeMetadataInput` and `ZarrV3ConsolidatedMetadataInput` resolve as annotations at run time, as a caller's `get_type_hints` reads them: types, not strings.""" + + def take(entry: ZarrV3NodeMetadataInput, member: ZarrV3ConsolidatedMetadataInput) -> None: + pass + + hints = get_type_hints(take) + assert {"entry", "member"} <= set(hints) + assert not isinstance(hints["entry"], str) + + +def test_error_a_conflict_is_reported_before_what_the_re_read_finds() -> None: + """A model read with the core `bytes` and an `endian`, given to a group whose private `bytes` takes none, is refused for the conflict first, and for the key that `bytes` does not take after it: the cause before its symptom.""" + child = ZarrV3ArrayMetadata( + {**ARRAY, "codecs": ({"name": "bytes", "configuration": {"endian": "little"}},)} + ) + with pytest.raises(MetadataValidationError) as raised: + _group(PRIVATE, a=child) + assert [problem.loc for problem in raised.value.problems] == [ + ("consolidated_metadata", "metadata", "a", "codecs", 0), + ("consolidated_metadata", "metadata", "a", "codecs", 0, "configuration", "endian"), + ] diff --git a/packages/zarr-metadata/tests/model/test_construction.py b/packages/zarr-metadata/tests/model/test_construction.py new file mode 100644 index 0000000000..185b6740cb --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_construction.py @@ -0,0 +1,310 @@ +"""A model checks itself when it is built, as pydantic's `__init__` does. + +A v3 model is built only by reading its document, and changed only by +`update`, which reads the document it makes: a document with a problem is +refused, with every problem, so no model is ever invalid. A v2 model, a +dataclass still, is refused at a `dataclasses.replace` into an invalid +one. `to_key_value` writes a model as it is, reading nothing. +""" + +from __future__ import annotations + +import dataclasses +import math +from collections import UserDict +from typing import TYPE_CHECKING, Any, cast + +import pytest + +from zarr_metadata.model import ( + MetadataValidationError, + ValidationProblem, + ZarrV2ArrayMetadata, + ZarrV2ConsolidatedMetadata, + ZarrV2GroupMetadata, + ZarrV3ArrayMetadata, + ZarrV3ConsolidatedMetadata, + ZarrV3GroupMetadata, +) +from zarr_metadata.v3.data_type.int8 import INT8_DATA_TYPE +from zarr_metadata.v3.definition import ( + CORE_AND_EXTENSIONS, +) + +if TYPE_CHECKING: + from collections.abc import Iterator + + from zarr_metadata.v3.definition import Nested + +ARRAY = ZarrV3ArrayMetadata.create_default( + shape=(4,), + attributes={"a": [1, None]}, + dimension_names=("x",), + acme={"must_understand": False}, +) +GROUP = ZarrV3GroupMetadata.from_json( + { + "zarr_format": 3, + "node_type": "group", + "attributes": {"a": 1}, + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": { + "x": ARRAY.to_json(), + "y": {"zarr_format": 3, "node_type": "group"}, + "y/z": ARRAY.to_json(), + }, + }, + } +) +V2_ARRAY = ZarrV2ArrayMetadata.create_default(shape=(4,), attributes={"a": 1}) +V2_GROUP = ZarrV2GroupMetadata.create_default(attributes={}) +V2_CONSOLIDATED = ZarrV2ConsolidatedMetadata.from_json( + {"zarr_consolidated_format": 1, "metadata": {"a/.zattrs": {"x": 1}}} +) + + +def _rebuilt(model: object, changes: dict[str, object]) -> Any: # noqa: ANN401 - the tests give the model as object + """`model`, a v2 model, built again of its document with `changes` in place of members, in its own scope: what `update` does, for the consolidated model too.""" + held = cast("Any", model) + return type(held)({**held.to_json(), **changes}, context=held.context) + + +_INLINE: dict[str, Any] = {"kind": "inline", "must_understand": False} + + +@pytest.mark.parametrize( + "model", + [ARRAY, GROUP, GROUP.consolidated_metadata], + ids=["array", "group", "consolidated"], +) +def test_a_v3_model_is_built_of_its_document_in_its_scope(model: object) -> None: + """A v3 model is its document read in its scope: built again of the two, it is the same model.""" + held = cast("Any", model) + assert type(held)(held.to_json(), context=held.context) == model + + +@pytest.mark.parametrize( + "model", + [ + V2_ARRAY, + V2_GROUP, + V2_CONSOLIDATED, + ], + ids=["v2-array", "v2-group", "v2-consolidated"], +) +def test_a_v2_model_is_built_as_a_read_builds_it(model: object) -> None: + """A v2 model is built of its document in its scope, as a read builds it: the constructor given the model's own document and scope builds an equal model.""" + held = cast("Any", model) + assert type(held)(held.to_json(), context=held.context) == model + + +@pytest.mark.parametrize( + ("model", "changes"), + [ + (ARRAY, {"shape": [4], "dimension_names": ["x"]}), + (ARRAY, {"attributes": UserDict({"a": [1, None]})}), + (GROUP, {"attributes": UserDict({"a": 1})}), + ], + ids=["array-lists", "array-mapping", "group-mapping"], +) +def test_a_v3_model_holds_its_members_as_a_read_refines_them( + model: object, changes: dict[str, object] +) -> None: + """Arrays as tuples and objects as dicts, as the read holds them, so a model updated with other containers is the model a read builds, and round-trips through its store.""" + changed = cast("Any", model).update(**changes) + assert changed == model + assert type(changed).from_key_value(changed.to_key_value()) == changed + + +@pytest.mark.parametrize( + ("model", "changes"), + [ + (V2_ARRAY, {"shape": range(4, 5), "chunks": [4], "attributes": {"a": 1}}), + (V2_CONSOLIDATED, {"metadata": {"a/.zattrs": UserDict({"x": 1})}}), + ], + ids=["v2-sequences", "v2-consolidated-mapping"], +) +def test_a_v2_model_holds_its_members_as_a_read_refines_them( + model: object, changes: dict[str, object] +) -> None: + """Arrays as tuples and objects as dicts, as the read holds them, so a v2 model built again with other containers is the model a read builds.""" + changed = _rebuilt(model, changes) + assert changed == model + assert type(changed).from_key_value(changed.to_key_value()) == changed + + +@pytest.mark.parametrize( + ("model", "member"), + [(ARRAY, "acme.x"), (GROUP, "attributes")], + ids=["array-extra-field", "group-attributes"], +) +def test_a_v3_model_shares_no_container_with_what_it_was_built_of( + model: object, member: str +) -> None: + """A model holds copies of the containers `update` is given: changing them afterwards changes nothing it holds.""" + held: dict[str, object] = {"must_understand": False, "z": {"y": 1}} + built = cast("Any", model).update(**{member: held}) + held["w"] = math.nan + cast("dict[str, object]", held["z"])["y"] = math.nan + expected = {"must_understand": False, "z": {"y": 1}} + if member == "attributes": + assert built.attributes == expected + else: + assert built.extra_fields[member] == expected + assert type(built).from_key_value(built.to_key_value()) == built + + +@pytest.mark.parametrize( + ("model", "member"), + [ + (V2_ARRAY, "attributes"), + (V2_GROUP, "attributes"), + (V2_CONSOLIDATED, "metadata"), + ], + ids=["v2-array-attributes", "v2-group-attributes", "v2-consolidated"], +) +def test_a_v2_model_shares_no_container_with_what_it_was_built_of( + model: object, member: str +) -> None: + """A v2 model holds copies of the containers it is built of.""" + held: dict[str, object] = {"acme.x": {"must_understand": False}} + built = _rebuilt(model, {member: held}) + held["acme.y"] = math.nan + cast("dict[str, object]", held["acme.x"])["z"] = math.nan + assert getattr(built, member) == {"acme.x": {"must_understand": False}} + assert type(built).from_key_value(built.to_key_value()) == built + + +def test_a_model_is_read_when_built_and_written_as_it_is() -> None: + """A model is read once, when it is built; `to_key_value` reads nothing; `update` reads the document it makes; a group reads the documents it holds once, as part of its own read.""" + values: list[object] = [] + + def counted( + configuration: object, nested: Nested, value: object + ) -> Iterator[ValidationProblem]: + values.append(value) + yield from () + + counting = dataclasses.replace(INT8_DATA_TYPE, fill_value_rules=counted) + scope = CORE_AND_EXTENSIONS.extended_with(counting) + model = ZarrV3ArrayMetadata.create_default(context=scope, data_type="int8", fill_value=3) + assert values == [3] + model.to_key_value() + assert values == [3] + changed = model.update(fill_value=4) + assert values == [3, 4] + group = ZarrV3GroupMetadata.create_default( + context=scope, consolidated_metadata={**_INLINE, "metadata": {"a": changed.to_json()}} + ) + assert values == [3, 4, 4] + group.to_key_value() + assert values == [3, 4, 4] + + +@pytest.mark.parametrize( + ("model", "changes", "problems"), + [ + (ARRAY, {"fill_value": math.nan}, [(("fill_value",), "invalid_value")]), + (ARRAY, {"dimension_names": ("x", "y")}, [(("dimension_names",), "invalid_value")]), + # Every problem: the grid and the names are each for one dimension. + ( + ARRAY, + {"shape": (4, 4)}, + [ + (("chunk_grid", "configuration", "chunk_shape"), "invalid_value"), + (("dimension_names",), "invalid_value"), + ], + ), + (ARRAY, {"attributes": {1: "a"}}, [(("attributes",), "invalid_type")]), + # Empty or not, a value that is no object is judged as one. + (ARRAY, {"attributes": []}, [(("attributes",), "invalid_type")]), + (ARRAY, {"attributes": None}, [(("attributes",), "invalid_type")]), + (GROUP, {"attributes": {1: "a"}}, [(("attributes",), "invalid_type")]), + (GROUP, {"attributes": ()}, [(("attributes",), "invalid_type")]), + (GROUP, {"acme": math.nan}, [(("acme",), "invalid_value")]), + ], + ids=[ + "fill-value-not-json", + "names-for-another-rank", + "shape-without-its-grid", + "attribute-key", + "attributes-empty-and-no-object", + "attributes-null", + "group-attribute-key", + "group-attributes-empty-and-no-object", + "group-extension-not-json", + ], +) +def test_error_a_v3_model_updated_into_an_invalid_one_is_refused_at_the_change( + model: object, changes: dict[str, object], problems: list[tuple[tuple[str | int, ...], str]] +) -> None: + """`update` reads the document it makes, so a change that makes an invalid one is refused with every problem: no model is invalid, however it came to be.""" + with pytest.raises(MetadataValidationError) as raised: + cast("Any", model).update(**changes) + assert [(found.loc, found.kind) for found in raised.value.problems] == problems + + +def test_error_consolidated_metadata_of_documents_at_bad_paths_is_refused() -> None: + """The consolidated member read on its own refuses documents at paths no node has, and reports a group missing above one, as the group's read does.""" + documents = { + "x": ARRAY.to_json(), + "x/a": ARRAY.to_json(), + "__b": ARRAY.to_json(), + "c/d": ARRAY.to_json(), + } + with pytest.raises(MetadataValidationError) as raised: + ZarrV3ConsolidatedMetadata({**_INLINE, "metadata": documents}) + assert [(found.loc, found.kind) for found in raised.value.problems] == [ + (("metadata", "__b"), "invalid_value"), + (("metadata", "x/a"), "invalid_value"), + (("metadata", "c"), "missing_key"), + ] + + +@pytest.mark.parametrize( + ("model", "changes", "problems"), + [ + (V2_ARRAY, {"order": "Q"}, [(("order",), "invalid_value")]), + (V2_ARRAY, {"chunks": (4, 4)}, [(("chunks",), "invalid_value")]), + (V2_GROUP, {"attributes": {1: "a"}}, [(("attributes",), "invalid_type")]), + ( + V2_CONSOLIDATED, + {"metadata": {"a/.zarray": {"x": math.nan}}}, + [(("metadata", "a/.zarray", "x"), "invalid_value")], + ), + ], + ids=[ + "v2-order", + "v2-chunks-for-another-rank", + "v2-group-attribute-key", + "v2-consolidated-entry-not-json", + ], +) +def test_error_a_v2_model_changed_by_hand_into_an_invalid_one_is_refused_at_the_change( + model: object, changes: dict[str, object], problems: list[tuple[tuple[str | int, ...], str]] +) -> None: + """As a read reads its document: no v2 model is invalid, however it came to be.""" + with pytest.raises(MetadataValidationError) as raised: + _rebuilt(model, changes) + assert [(found.loc, found.kind) for found in raised.value.problems] == problems + + +def test_error_v2_create_default_refuses_chunks_its_default_shape_does_not_take() -> None: + # Overriding chunks without shape keeps the scalar default shape (), as + # the v3 model keeps its default shape, and refuses a grid it does not fit. + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ArrayMetadata.create_default(chunks=(10, 10)) + assert [(found.loc, found.kind) for found in raised.value.problems] == [ + (("chunks",), "invalid_value") + ] + + +def test_error_consolidated_metadata_paths_are_strings() -> None: + """A `metadata` member keyed by what is no string is a problem of the member, as the read reports it, not a `TypeError`.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV3ConsolidatedMetadata({**_INLINE, "metadata": {1: ARRAY.to_json()}}) + assert [(found.loc, found.kind) for found in raised.value.problems] == [ + (("metadata",), "invalid_type") + ] diff --git a/packages/zarr-metadata/tests/model/test_extension_points.py b/packages/zarr-metadata/tests/model/test_extension_points.py index 1459e4ec48..bfd52df7b0 100644 --- a/packages/zarr-metadata/tests/model/test_extension_points.py +++ b/packages/zarr-metadata/tests/model/test_extension_points.py @@ -13,6 +13,7 @@ from typing import Any, cast import pytest +from typing_extensions import TypedDict from zarr_metadata._json import arrays_to_tuples from zarr_metadata.model import ( @@ -24,8 +25,12 @@ validate_array_metadata_v3, validate_group_metadata_v3, ) -from zarr_metadata.v3.codec.gzip import GZIP_CODEC -from zarr_metadata.v3.definition import CORE, CORE_AND_EXTENSIONS, CodecDefinition, Context +from zarr_metadata.v3.definition import ( + CORE, + CORE_AND_EXTENSIONS, + CodecDefinition, + Context, +) BYTES = {"name": "bytes", "configuration": {"endian": "little"}} @@ -36,12 +41,17 @@ def _document(**fields: object) -> dict[str, Any]: return cast("dict[str, Any]", arrays_to_tuples(document)) +class LenientGzipConfiguration(TypedDict, closed=True): + """A gzip configuration whose `level` is any integer: the bound is the type's, so taking any is a type of its own.""" + + level: int + + LENIENT_GZIP = CodecDefinition( name="gzip", - configuration=GZIP_CODEC.configuration, + configuration=LenientGzipConfiguration, kind="bytes_bytes", size="dynamic", - rules=lambda configuration, nested: [], ) """A reader's own gzip, which takes any level: a scope can grow, and substitute.""" @@ -78,13 +88,13 @@ def test_a_document_reads_through_the_definitions_in_its_scope( document: dict[str, Any], context: Context ) -> None: # What the scope holds judges; what it does not is left to the reader. - # The model reads and writes in the same scope as the validators. + # The model reads in the same scope as the validators, and what it + # writes reads back in that scope as the same model. assert validate_array_metadata_v3(document, context=context) == () assert is_array_metadata_v3(document, context=context) assert parse_array_metadata_v3(document, context=context) is not None model = ZarrV3ArrayMetadata.from_json(document, context=context) - written = model.to_key_value(context=context) - assert ZarrV3ArrayMetadata.from_key_value(written, context=context) == model + assert ZarrV3ArrayMetadata.from_key_value(model.to_key_value(), context=context) == model @pytest.mark.parametrize( @@ -179,20 +189,7 @@ def test_an_array_in_a_group_s_consolidated_metadata_reads_in_the_group_s_scope( lenient = CORE.extended_with(LENIENT_GZIP) assert validate_group_metadata_v3(group, context=lenient) == () model = ZarrV3GroupMetadata.from_json(group, context=lenient) - written = model.to_key_value(context=lenient) - assert ZarrV3GroupMetadata.from_key_value(written, context=lenient) == model - - -def test_error_a_model_is_not_written_in_a_scope_that_refuses_it() -> None: - # The writer judges in the scope it is given, as the reader does: the - # default one refuses the level the reader's own gzip took. - document = _document(codecs=[BYTES, {"name": "gzip", "configuration": {"level": 99}}]) - model = ZarrV3ArrayMetadata.from_json(document, context=CORE.extended_with(LENIENT_GZIP)) - with pytest.raises(MetadataValidationError) as raised: - model.to_key_value() - assert [(problem.loc, problem.kind) for problem in raised.value.problems] == [ - (("codecs", 1, "configuration", "level"), "invalid_value") - ] + assert ZarrV3GroupMetadata.from_key_value(model.to_key_value(), context=lenient) == model def test_error_an_envelope_is_judged_once() -> None: diff --git a/packages/zarr-metadata/tests/model/test_group.py b/packages/zarr-metadata/tests/model/test_group.py index 845e303fd4..d840574b5d 100644 --- a/packages/zarr-metadata/tests/model/test_group.py +++ b/packages/zarr-metadata/tests/model/test_group.py @@ -3,40 +3,71 @@ import copy import dataclasses import json +import pickle from collections import UserDict -from collections.abc import Callable +from collections.abc import Callable, Iterator +from typing import Any, cast import pytest from tests.model._cases import mutate_nested_containers +from zarr_metadata._common import JSONValue, ZarrV3NamedConfigJSON from zarr_metadata._json import ( + JSON_DEPTH, MetadataValidationError, ValidationProblem, arrays_to_tuples, + json_text, ) -from zarr_metadata.model import UNSET -from zarr_metadata.model._array import ZarrV3ArrayMetadata +from zarr_metadata.model import ( + UNSET, + is_array_metadata_v3, +) +from zarr_metadata.model._array import ZarrV3ArrayMetadata, ZarrV3ArrayMetadataUpdate from zarr_metadata.model._group import ( ZarrV2ConsolidatedMetadata, ZarrV2GroupMetadata, - ZarrV2GroupMetadataPartial, + ZarrV2GroupMetadataUpdate, ZarrV3ConsolidatedMetadata, ZarrV3GroupMetadata, - ZarrV3GroupMetadataPartial, + ZarrV3GroupMetadataReading, + ZarrV3GroupMetadataUpdate, + ZarrV3UnknownNodeReading, + is_group_metadata_v3, + node_metadata_from_json_v3, + node_metadata_from_key_value_v3, + parse_group_metadata_v3, + read_group_metadata_v3, + read_node_metadata_v3, + validate_group_metadata_v3, + validate_node_metadata_v3, ) from zarr_metadata.model._validation import ( + ZarrV3ArrayMetadataReading, is_group_metadata_v2, - is_group_metadata_v3, parse_group_metadata_v2, - parse_group_metadata_v3, validate_group_metadata_v2, - validate_group_metadata_v3, ) from zarr_metadata.v2.group import ( ZarrV2GroupMetadataJSON, ZarrV2GroupMetadataJSONPartial, ZarrV2ZGroupJSON, ) +from zarr_metadata.v3.array import ZarrV3ArrayMetadataJSONPartial +from zarr_metadata.v3.codec.gzip import GZIP_CODEC, GzipCodecConfiguration +from zarr_metadata.v3.definition import ( + CORE, + CORE_AND_EXTENSIONS, + CodecDefinition, + EmptyConfiguration, + Nested, + RefusedField, + UnclaimedField, + resolve, +) +from zarr_metadata.v3.group import ZarrV3GroupMetadataJSONPartial + +_INLINE_ENVELOPE: dict[str, Any] = {"kind": "inline", "must_understand": False} # --- ZarrV3GroupMetadata --------------------------------------------------- @@ -81,26 +112,6 @@ def test_group_v3_json_extra_field_roundtrips_as_must_understand() -> None: assert model.must_understand_fields == {"ext": (1, 2)} -def test_group_v3_extra_fields_overlap_rejected() -> None: - """Constructing a v3 group model with extra_fields shadowing a standard key raises.""" - with pytest.raises(ValueError, match="Extra fields"): - ZarrV3GroupMetadata( - attributes={}, - consolidated_metadata=UNSET, - extra_fields={"node_type": {"name": "x", "must_understand": False}}, - ) - - -def test_group_v3_consolidated_extra_field_rejected() -> None: - """extra_fields may not shadow the consolidated_metadata convention key.""" - with pytest.raises(ValueError, match="Extra fields"): - ZarrV3GroupMetadata( - attributes={}, - consolidated_metadata=UNSET, - extra_fields={"consolidated_metadata": {"name": "x", "must_understand": False}}, - ) - - def test_group_v3_missing_required_key() -> None: """parse_group_metadata_v3 reports each missing required key.""" with pytest.raises(MetadataValidationError, match="node_type"): @@ -135,11 +146,58 @@ def test_group_zarr_format_rejects_float( assert [(p.loc, p.kind) for p in validate(document)] == [(("zarr_format",), "invalid_value")] -def test_group_v2_rejects_unknown_document_member() -> None: - """The closed v2 merged-document shape rejects undeclared members.""" - assert [(p.loc, p.kind) for p in validate_group_metadata_v2({"zarr_format": 2, "x": 1})] == [ - (("x",), "invalid_value") - ] +def _v2_consolidated_problems(document: object) -> list[tuple[tuple[str | int, ...], str]]: + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ConsolidatedMetadata.from_json(document) + return [(p.loc, p.kind) for p in raised.value.problems] + + +def _v3_field_problems(field: object) -> list[tuple[tuple[str | int, ...], str]]: + return [(p.loc, p.kind) for p in resolve(field, CodecDefinition, CORE_AND_EXTENSIONS)[1]] + + +@pytest.mark.parametrize( + ("problems", "expected"), + [ + ( + lambda: [ + (p.loc, p.kind) for p in validate_group_metadata_v2({"zarr_format": 2, "x": 1}) + ], + [(("x",), "unknown_key")], + ), + ( + lambda: _v2_consolidated_problems( + {"zarr_consolidated_format": 1, "metadata": {}, "x": 1} + ), + [(("x",), "unknown_key")], + ), + ( + lambda: [ + (p.loc, p.kind) + for p in validate_group_metadata_v3( + _group(consolidated_metadata={**_inline(), "x": 1}) + ) + ], + [(("consolidated_metadata", "x"), "unknown_key")], + ), + ( + lambda: _v3_field_problems({"name": "gzip", "configuration": {"level": 1}, "x": 1}), + [(("x",), "unknown_key")], + ), + ( + lambda: _v3_field_problems({"name": "gzip", "configuration": {"level": 1, "x": 1}}), + [(("configuration", "x"), "unknown_key")], + ), + ], + ids=["v2-group", "v2-consolidated", "v3-consolidated", "v3-field", "v3-configuration"], +) +def test_a_member_a_closed_object_does_not_declare_is_an_unknown_key( + problems: Callable[[], list[tuple[tuple[str | int, ...], str]]], + expected: list[tuple[tuple[str | int, ...], str]], +) -> None: + # Wherever it sits, so a reader that tolerates what another writer + # added -- NCZarr's `_nczarr_*` keys -- filters by kind. + assert problems() == expected @pytest.mark.parametrize( @@ -214,6 +272,210 @@ def test_group_v3_update() -> None: # --- ZarrV2GroupMetadata --------------------------------------------------- +def test_a_v2_group_nested_as_deep_as_a_reader_walks_is_read_and_written() -> None: + """The v2 group copied with `copy.deepcopy`, two frames a level, and overflowed on documents its validator accepts.""" + attributes: dict[str, object] = {} + for _ in range(JSON_DEPTH - 2): + attributes = {"x": attributes} + document = {"zarr_format": 2, "attributes": attributes} + assert validate_group_metadata_v2(document) == () + model = ZarrV2GroupMetadata.from_json(document) + assert json_text(model.to_json()) == json_text(cast("JSONValue", document)) + assert ZarrV2GroupMetadata.from_key_value(model.to_key_value()) == model + assert pickle.loads(pickle.dumps(model)) == model + assert copy.deepcopy(model) == model + problems = validate_group_metadata_v2({**document, "attributes": {"x": attributes}}) + assert [(problem.kind, len(problem.loc)) for problem in problems] == [ + ("invalid_value", JSON_DEPTH) + ] + + +def _chain_of_groups(documents: int) -> dict[str, object]: + """A group holding, in its consolidated metadata, a group holding a group..., `documents` of them below the root, each listing only the next: a document sits three levels below the one holding it.""" + document: dict[str, object] = {"zarr_format": 3, "node_type": "group"} + for _ in range(documents): + document = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": {"a": document}, + }, + } + return document + + +def test_a_chain_of_consolidated_groups_is_bounded_by_the_levels_a_reader_walks() -> None: + """Each document is read from where it sits in the one handed in, so a chain is bounded as any nesting is: the reader took three frames a document, unbounded, and a 20 KB chain overflowed.""" + depth = f"nested deeper than the {JSON_DEPTH} levels a reader walks" + deepest = (JSON_DEPTH - 1) // 3 + problems = validate_group_metadata_v3(_chain_of_groups(deepest)) + assert all(len(problem.loc) < JSON_DEPTH for problem in problems) + assert {problem.message for problem in problems} == { + 'expected a node the group lists, got "/a/a", which "/a" lists alone' + } + for documents in (deepest + 1, 350): + problems = validate_group_metadata_v3(_chain_of_groups(documents)) + assert [ + (len(problem.loc), problem.message) + for problem in problems + if len(problem.loc) >= JSON_DEPTH + ] == [(JSON_DEPTH, depth)] + with pytest.raises(MetadataValidationError): + ZarrV3GroupMetadata.from_json(_chain_of_groups(documents)) + + +def test_a_consolidated_document_is_read_from_where_it_sits() -> None: + """Its levels are counted from the root of the document handed in, so what `read_node_metadata_v3` admits alone can sit too deep inside; its own reading locates the problem in it.""" + depth = f"nested deeper than the {JSON_DEPTH} levels a reader walks" + + def nested(levels: int) -> dict[str, object]: + value: dict[str, object] = {} + for _ in range(levels): + value = {"x": value} + return value + + # A document sits three levels down, its attribute, or a member of an + # extra field, two more. + group = { + "zarr_format": 3, + "node_type": "group", + "attributes": {"a": nested(JSON_DEPTH - 4)}, + "acme_extra": {"must_understand": False, "a": nested(JSON_DEPTH - 4)}, + } + array = { + **ZarrV3ArrayMetadata.create_default(shape=(2,)).to_json(), + "data_type": "acme.deep", + "fill_value": {"a": nested(JSON_DEPTH - 4)}, + "codecs": [ + {"name": "bytes", "configuration": {"endian": "little"}}, + {"name": "acme.x", "configuration": {"y": nested(JSON_DEPTH - 6)}}, + ], + } + assert validate_node_metadata_v3(group) == () + assert validate_node_metadata_v3(array) == () + document = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": {"g": group, "h": array}, + }, + } + reading = read_group_metadata_v3(document) + assert [(len(problem.loc), problem.message) for problem in reading.problems] == [ + (JSON_DEPTH, depth) + ] * 4 + assert [ + (problem.loc[:2], len(problem.loc)) for problem in reading.consolidated["g"].problems + ] == [(("acme_extra", "a"), JSON_DEPTH - 3), (("attributes", "a"), JSON_DEPTH - 3)] + assert [ + (problem.loc[:2], len(problem.loc)) for problem in reading.consolidated["h"].problems + ] == [ + (("fill_value", "a"), JSON_DEPTH - 3), + (("codecs", 1), JSON_DEPTH - 3), + ] + + +def test_a_document_a_level_before_the_cap_holds_its_scalars_and_nothing_else() -> None: + """Each container member sits past the cap, and is judged where it sits before it is walked, as `refine_json` of the whole would judge it; the scalars are read.""" + depth = f"nested deeper than the {JSON_DEPTH} levels a reader walks" + + def a_level_before_the_cap(document: dict[str, object]) -> dict[str, object]: + for _ in range((JSON_DEPTH - 1) // 3): + document = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": {"a": document}, + }, + } + return document + + scalars: dict[str, object] = {"zarr_format": 3, "node_type": "group"} + problems = validate_group_metadata_v3(a_level_before_the_cap(scalars)) + assert all(len(problem.loc) < JSON_DEPTH for problem in problems) + array = ZarrV3ArrayMetadata.create_default(shape=(2,)).to_json() + for document in ({**scalars, "attributes": {"a": 1}}, dict(array)): + containers = sorted( + key for key, item in document.items() if isinstance(item, (dict, list, tuple)) + ) + assert len(containers) != 0 + problems = validate_group_metadata_v3(a_level_before_the_cap(document)) + assert sorted( + (problem.loc[-1], len(problem.loc), problem.message) + for problem in problems + if len(problem.loc) >= JSON_DEPTH + ) == [(key, JSON_DEPTH, depth) for key in containers] + + +def test_a_listing_key_too_long_to_write_is_shown_by_its_size( + interpreter_writes_4300_digits: None, +) -> None: + document = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": {10**5000: 1}, + }, + } + (problem,) = validate_group_metadata_v3(document) + assert problem.message == f"non-string key an integer of {(10**5000).bit_length()} bits" + + +def test_a_group_reads_the_documents_it_holds_from_where_they_sit() -> None: + """A child valid alone, nested to the last level a reader walks, sits three levels deeper as a document a group holds: the group's read refuses it at the cap, as the validator does, and so does the consolidated member read on its own, from where it sits under the group's key.""" + # The innermost object sits at the last level a reader walks, alone; + # three deeper as a document a group holds. + nested: dict[str, object] = {} + for _ in range(JSON_DEPTH - 3): + nested = {"x": nested} + child = ZarrV3GroupMetadata.from_json( + {"zarr_format": 3, "node_type": "group", "attributes": {"a": nested}} + ) + written = { + **ZarrV3GroupMetadata.create_default().to_json(), + "consolidated_metadata": {**_INLINE_ENVELOPE, "metadata": {"a": child.to_json()}}, + } + depth = f"nested deeper than the {JSON_DEPTH} levels a reader walks" + assert [(len(p.loc), p.message) for p in validate_group_metadata_v3(written)] == [ + (JSON_DEPTH, depth) + ] + with pytest.raises(MetadataValidationError): + ZarrV3GroupMetadata.from_json(written) + with pytest.raises(MetadataValidationError): + ZarrV3ConsolidatedMetadata.from_json(written["consolidated_metadata"]) + + +def test_a_v2_consolidated_document_is_read_to_the_levels_a_reader_walks() -> None: + """An entry past the cap is the depth problem, not `RecursionError`: the document was normalized whole, a frame a level without bound, before any entry was refined.""" + + def nested(levels: int) -> dict[str, object]: + value: dict[str, object] = {} + for _ in range(levels): + value = {"x": value} + return value + + # An entry sits two levels down; its innermost object at the last + # level a reader walks. + document = {"zarr_consolidated_format": 1, "metadata": {"a/.zattrs": nested(JSON_DEPTH - 3)}} + model = ZarrV2ConsolidatedMetadata.from_json(document) + assert json_text(model.to_json()) == json_text(cast("JSONValue", document)) + for levels in (JSON_DEPTH - 2, 2000): + deeper = {**document, "metadata": {"a/.zattrs": nested(levels)}} + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ConsolidatedMetadata.from_json(deeper) + assert [(len(p.loc), p.kind) for p in raised.value.problems] == [ + (JSON_DEPTH, "invalid_value") + ] + + def test_group_v2_key_value_split() -> None: """v2 to_key_value writes .zgroup and .zattrs; from_key_value merges them.""" model = ZarrV2GroupMetadata.create_default(attributes={"a": 1}) @@ -232,7 +494,7 @@ def test_v2_group_from_key_value_rejects_zgroup_extra_members(extra_key: str) -> ZarrV2GroupMetadata.from_key_value({".zgroup": json.dumps(doc).encode()}) assert [(problem.loc, problem.kind) for problem in exc_info.value.problems] == [ - ((extra_key,), "invalid_value") + ((extra_key,), "unknown_key") ] @@ -265,7 +527,7 @@ def test_group_v2_omits_empty_attributes() -> None: def test_group_v2_not_a_mapping() -> None: """parse_group_metadata_v2 rejects a non-mapping document.""" - with pytest.raises(MetadataValidationError, match="expected a mapping"): + with pytest.raises(MetadataValidationError, match="expected an object"): parse_group_metadata_v2([1, 2, 3]) @@ -279,17 +541,22 @@ def test_group_v2_missing_required_key() -> None: def test_group_partial_keys_match_settable_model_fields() -> None: - """Each group partial TypedDict must list exactly the settable model fields. + """The v2 group partial TypedDict lists exactly the settable model fields. - Guards against drift: adding/removing a settable field on a group model + Guards against drift: adding/removing a settable field on the model without updating its `*Partial` TypedDict fails here. """ - for model_cls, partial_cls in ( - (ZarrV3GroupMetadata, ZarrV3GroupMetadataPartial), - (ZarrV2GroupMetadata, ZarrV2GroupMetadataPartial), + assert set(ZarrV2GroupMetadataUpdate.__annotations__) == {"attributes"} + + +def test_update_takes_every_member_of_the_document_it_may_change() -> None: + """Each v3 model's `update` takes each member of its document but `zarr_format` and `node_type`, which it cannot change; a group's declares `consolidated_metadata` too, a member the document leaves to its extra items, since `update` takes node models there.""" + fixed = {"zarr_format", "node_type"} + for update, partial, declared in ( + (ZarrV3ArrayMetadataUpdate, ZarrV3ArrayMetadataJSONPartial, set[str]()), + (ZarrV3GroupMetadataUpdate, ZarrV3GroupMetadataJSONPartial, {"consolidated_metadata"}), ): - settable = {f.name for f in dataclasses.fields(model_cls) if f.init} - assert set(partial_cls.__annotations__) == settable + assert set(update.__annotations__) == (set(partial.__annotations__) - fixed) | declared # --- ZarrV3ConsolidatedMetadata -------------------------------------------- @@ -315,10 +582,10 @@ def test_consolidated_v3_roundtrip() -> None: assert model.to_json() == doc -def test_consolidated_v3_must_understand_true_rejected() -> None: - """ZarrV3ConsolidatedMetadata enforces must_understand=False at runtime.""" - with pytest.raises(ValueError, match="must_understand"): - ZarrV3ConsolidatedMetadata(must_understand=True, metadata={}) +def test_error_consolidated_v3_must_understand_is_not_a_member_to_set() -> None: + """`must_understand` is `False` by declaration, as `kind` is `"inline"`: not a member the constructor takes.""" + with pytest.raises(TypeError, match="must_understand"): + cast("Any", ZarrV3ConsolidatedMetadata)(must_understand=True, metadata={}) def test_consolidated_v3_from_json_must_understand_true_rejected() -> None: @@ -339,10 +606,491 @@ def test_consolidated_v3_entry_without_node_type_rejected() -> None: def test_consolidated_v3_not_a_mapping() -> None: """from_json rejects a non-mapping consolidated document.""" - with pytest.raises(MetadataValidationError, match="expected a mapping"): + with pytest.raises(MetadataValidationError, match="expected an object"): ZarrV3ConsolidatedMetadata.from_json(5) +# --- read_group_metadata_v3 ------------------------------------------------ + +LITTLE: ZarrV3NamedConfigJSON = {"name": "bytes", "configuration": {"endian": "little"}} +GZIP_99 = {"name": "gzip", "configuration": {"level": 99}} + +A: tuple[str, ...] = ("consolidated_metadata", "metadata", "a") +"""Where the document at path `a` in a group's consolidated metadata sits in the group's.""" + + +def _array(**members: object) -> dict[str, object]: + return {**ZarrV3ArrayMetadata.create_default(shape=(4,)).to_json(), **members} + + +def _inline(**documents: object) -> dict[str, object]: + return {"kind": "inline", "must_understand": False, "metadata": documents} + + +def _group(**members: object) -> dict[str, object]: + return {"zarr_format": 3, "node_type": "group", **members} + + +def test_error_a_consolidated_envelope_reports_the_members_it_lacks_in_its_order() -> None: + document = _group(consolidated_metadata={"kind": "inline"}) + assert [(p.loc, p.kind) for p in validate_group_metadata_v3(document)] == [ + (("consolidated_metadata", "must_understand"), "missing_key"), + (("consolidated_metadata", "metadata"), "missing_key"), + ] + + +ACME_X = CodecDefinition( + name="acme.x", configuration=EmptyConfiguration, kind="bytes_bytes", size="dynamic" +) + + +def test_group_update_keeps_the_documents_it_holds() -> None: + """Read in no scope again: one read in a scope the call's does not claim keeps each field as it was read.""" + scope = CORE_AND_EXTENSIONS.extended_with(ACME_X) + child = ZarrV3ArrayMetadata.create_default( + context=scope, shape=(4,), codecs=(LITTLE, {"name": "acme.x"}) + ) + group = ZarrV3GroupMetadata.create_default( + context=scope, + consolidated_metadata={**_INLINE_ENVELOPE, "metadata": {"a": child.to_json()}}, + ) + updated = group.update(attributes={"k": 1}) + assert updated.attributes == {"k": 1} + assert updated.context == scope + assert updated.consolidated_metadata == group.consolidated_metadata + + +def test_group_update_reads_the_documents_it_is_given_in_its_scope() -> None: + """And `UNSET` leaves them out. `zstd` is an extension, which `CORE` leaves unclaimed.""" + zstd = {"name": "zstd", "configuration": {"level": 3, "checksum": False}} + member = cast("Any", _inline(a=_array(codecs=[LITTLE, zstd]))) + updated = ZarrV3GroupMetadata.create_default(context=CORE).update(consolidated_metadata=member) + assert updated.consolidated_metadata is not UNSET + child = updated.consolidated_metadata.metadata["a"] + assert isinstance(child, ZarrV3ArrayMetadata) + assert isinstance(child.codecs[1], UnclaimedField) + removed = updated.update(consolidated_metadata=UNSET) + assert removed.consolidated_metadata is UNSET + + +def _fields_of_an_array(*at: str | int) -> list[tuple[str | int, ...]]: + """Where a default array's fields sit, under `at`.""" + points = ("data_type", "chunk_grid", "chunk_key_encoding") + return [*((*at, point) for point in points), (*at, "codecs", 0)] + + +@pytest.mark.parametrize( + ("document", "paths", "locs"), + [ + (_group(), [], []), + ( + _group(consolidated_metadata=_inline(a=_array(), g=_group())), + ["a", "g"], + _fields_of_an_array(*A), + ), + # Every node below the group, each below a group, at its path. + ( + _group(consolidated_metadata=_inline(g=_group(), **{"g/b": _array()})), + ["g", "g/b"], + _fields_of_an_array("consolidated_metadata", "metadata", "g/b"), + ), + # A group's consolidated metadata in a group's, listing what the group + # lists too: each field located from the root of the outer document. + ( + _group( + consolidated_metadata=_inline( + g=_group(consolidated_metadata=_inline(b=_array())), **{"g/b": _array()} + ) + ), + ["g", "g/b"], + [ + *_fields_of_an_array( + "consolidated_metadata", + "metadata", + "g", + "consolidated_metadata", + "metadata", + "b", + ), + *_fields_of_an_array("consolidated_metadata", "metadata", "g/b"), + ], + ), + ], + ids=["no-consolidated-metadata", "an-array-and-a-group", "paths", "nested"], +) +def test_a_group_reads_each_document_its_consolidated_metadata_holds( + document: dict[str, object], paths: list[str], locs: list[tuple[str | int, ...]] +) -> None: + reading = read_group_metadata_v3(document) + assert reading.problems == () + assert list(reading.consolidated) == paths + assert [loc for loc, _ in reading.fields()] == locs + model = reading.metadata + assert model is not None + assert model == ZarrV3GroupMetadata.from_json(document) + # It holds the model each document's own reading built. + consolidated = model.consolidated_metadata + held = {} if consolidated is UNSET else consolidated.metadata + assert all(held[path] is reading.consolidated[path].metadata for path in paths) + + +@pytest.mark.parametrize( + ("path", "fault"), + [ + ("", "is the group's own"), + ("/a", 'starts with "/"'), + ("a/", 'ends with "/"'), + ("a//b", 'holds an empty name between two "/"'), + (".", 'holds ".", a name that is periods alone'), + ("a/../b", 'holds "..", a name that is periods alone'), + ("__a", 'holds "__a", a name that starts with the reserved "__"'), + ("zarr.json", 'holds "zarr.json", a name that is the reserved "zarr.json"'), + ], +) +def test_error_a_document_in_consolidated_metadata_is_at_a_node_s_path_below_the_group( + path: str, fault: str +) -> None: + # Its node names, joined by "/", as the reference implementation keeps + # it: the group's own path, "/", and it make the node's. + document = _group(consolidated_metadata=_inline(**{path: _group()})) + message = f"expected the path of a node below the group, got {json.dumps(path)}, which {fault}" + assert validate_group_metadata_v3(document) == ( + ValidationProblem(("consolidated_metadata", "metadata", path), message, "invalid_value"), + ) + + +@pytest.mark.parametrize( + ("documents", "path"), + [ + ({"a": _array(), "a/b": _group()}, "a/b"), + # However many groups are missing between them: none would help. + ({"a": _array(), "a/b/c": _array()}, "a/b/c"), + ], + ids=["child", "descendant"], +) +def test_error_no_document_in_consolidated_metadata_is_below_an_array( + documents: dict[str, object], path: str +) -> None: + # "Group nodes may have children but array nodes may not." A message + # names each node by its path in the hierarchy below the group. + document = _group(consolidated_metadata=_inline(**documents)) + message = f'expected a node below a group, got "/{path}", below the array "/a"' + assert validate_group_metadata_v3(document) == ( + ValidationProblem(("consolidated_metadata", "metadata", path), message, "invalid_value"), + ) + + +def test_a_document_of_no_node_type_is_taken_as_a_group_s() -> None: + # Its own problem is reported where it sits, and the nodes below it are + # not refused for it. + document = _group( + consolidated_metadata=_inline(a={"zarr_format": 3, "node_type": "x"}, **{"a/b": _array()}) + ) + assert [(p.loc, p.kind) for p in validate_group_metadata_v3(document)] == [ + (("consolidated_metadata", "metadata", "a", "node_type"), "invalid_value") + ] + + +def test_error_consolidated_metadata_holds_the_group_holding_each_document() -> None: + # The nearest group missing above a node, once, counting those above it + # up to the group holding the documents. + document = _group(consolidated_metadata=_inline(**{"a/b/c": _array(), "a/b/d": _array()})) + assert [(p.loc, p.message, p.kind) for p in validate_group_metadata_v3(document)] == [ + ( + ("consolidated_metadata", "metadata", "a/b"), + 'missing the group holding "/a/b/c", and 1 group above it', + "missing_key", + ), + ] + + +def test_error_a_consolidated_metadata_key_that_is_not_a_string() -> None: + listing: dict[object, object] = {"a": _array(), 1: _array()} + document = _group(consolidated_metadata={**_inline(), "metadata": listing}) + assert [(p.loc, p.kind) for p in validate_group_metadata_v3(document)] == [ + (("consolidated_metadata", "metadata"), "invalid_type") + ] + + +NESTED = ("consolidated_metadata", "metadata", "g", "consolidated_metadata", "metadata") +"""Where the own listing of the group at `g` sits in the outer document.""" + + +@pytest.mark.parametrize( + ("listed", "flat", "expected"), + [ + # What a listed group lists itself is what the group lists, too. + ({"b": _array()}, {"g/b": _array()}, []), + # A node the listed group lists alone would be dropped by the + # reference reader, which keeps the flat listing. + ( + {"b": _array()}, + {}, + [ + ( + (*NESTED, "b"), + 'expected a node the group lists, got "/g/b", which "/g" lists alone', + "invalid_value", + ) + ], + ), + # Nor may the two listings disagree on what a node is. + ( + {"b": _array()}, + {"g/b": _group()}, + [ + ( + (*NESTED, "b"), + 'expected a group, as the group lists "/g/b", got an array', + "invalid_value", + ) + ], + ), + # A deeper listing is the listed group's own to judge, when its + # document is read: its problem is that document's, at its place. + ( + {"h": _group(consolidated_metadata=_inline(x=_array()))}, + {"g/h": _group()}, + [ + ( + (*NESTED, "h", "consolidated_metadata", "metadata", "x"), + 'expected a node the group lists, got "/h/x", which "/h" lists alone', + "invalid_value", + ) + ], + ), + ], + ids=["agreeing", "listed-alone", "contradicting", "deeper"], +) +def test_a_listed_group_s_own_listing_lists_what_the_group_lists( + listed: dict[str, object], + flat: dict[str, object], + expected: list[tuple[tuple[str, ...], str, str]], +) -> None: + document = _group( + consolidated_metadata=_inline(g=_group(consolidated_metadata=_inline(**listed)), **flat) + ) + problems = validate_group_metadata_v3(document) + assert [(p.loc, p.message, p.kind) for p in problems] == expected + # The consolidated member read on its own refuses what the group's read + # reports: a listed group whose own listing is wrong is refused. + listing = {"g": _group(consolidated_metadata=_inline(**listed)), **flat} + deeper = [(loc, message, kind) for loc, message, kind in expected if len(loc) > len(NESTED) + 1] + if len(deeper) != 0: + with pytest.raises(MetadataValidationError) as inner: + node_metadata_from_json_v3(listing["g"]) + assert [(p.loc, p.message, p.kind) for p in inner.value.problems] == [ + (loc[3:], message, kind) for loc, message, kind in deeper + ] + return + if expected == []: + assert ZarrV3ConsolidatedMetadata(_inline(**listing)).metadata.keys() == listing.keys() + return + with pytest.raises(MetadataValidationError) as raised: + ZarrV3ConsolidatedMetadata(_inline(**listing)) + assert [(p.loc[1:], p.message, p.kind) for p in raised.value.problems] == [ + (loc[2:], message, kind) for loc, message, kind in expected + ] + + +def test_each_document_its_consolidated_metadata_holds_is_read_once() -> None: + reads: list[GzipCodecConfiguration] = [] + + def counted( + configuration: GzipCodecConfiguration, nested: Nested + ) -> Iterator[ValidationProblem]: + reads.append(configuration) + yield from () + + scope = CORE_AND_EXTENSIONS.extended_with(dataclasses.replace(GZIP_CODEC, rules=counted)) + array = _array(codecs=[LITTLE, {"name": "gzip", "configuration": {"level": 5}}]) + group = _group(consolidated_metadata=_inline(a=array)) + assert read_group_metadata_v3(group, context=scope).problems == () + assert reads == [{"level": 5}] + ZarrV3GroupMetadata.from_json(group, context=scope) + assert len(reads) == 2 + + +def test_error_a_group_document_that_is_not_an_object_reads_as_nothing() -> None: + not_an_object = ValidationProblem((), "expected an object", "invalid_type") + assert read_group_metadata_v3([1]) == ZarrV3GroupMetadataReading(problems=(not_an_object,)) + + +@pytest.mark.parametrize( + ("document", "problems", "models", "refused"), + [ + # The group's own member: each document it holds still has its model. + ( + _group(attributes=5, consolidated_metadata=_inline(a=_array())), + [(("attributes",), "invalid_type")], + {"a": True}, + [], + ), + # A document it holds: that one has none, and its sibling has one; + # the field refused is found where it sits, like every field. + ( + _group(consolidated_metadata=_inline(a=_array(codecs=[LITTLE, GZIP_99]), b=_array())), + [((*A, "codecs", 1, "configuration", "level"), "invalid_value")], + {"a": False, "b": True}, + [(*A, "codecs", 1)], + ), + # One of no node type is read as none, and nothing else of it is. + ( + _group(consolidated_metadata=_inline(a={"zarr_format": 3})), + [((*A, "node_type"), "missing_key")], + {"a": False}, + [], + ), + ], + ids=["group", "consolidated-document", "no-node-type"], +) +def test_error_a_group_document_with_a_problem_reads_as_no_model( + document: dict[str, object], + problems: list[tuple[tuple[str | int, ...], str]], + models: dict[str, bool], + refused: list[tuple[str | int, ...]], +) -> None: + reading = read_group_metadata_v3(document) + assert [(p.loc, p.kind) for p in reading.problems] == problems + assert reading.metadata is None + assert { + path: read.metadata is not None for path, read in reading.consolidated.items() + } == models + assert [loc for loc, field in reading.fields() if isinstance(field, RefusedField)] == refused + + +# --- read_node_metadata_v3 ------------------------------------------------- + + +@pytest.mark.parametrize( + ("document", "reading", "problems"), + [ + (_array(), ZarrV3ArrayMetadataReading, []), + (_group(attributes={"a": 1}), ZarrV3GroupMetadataReading, []), + (_array(fill_value=300), ZarrV3ArrayMetadataReading, [(("fill_value",), "invalid_value")]), + (_group(attributes=5), ZarrV3GroupMetadataReading, [(("attributes",), "invalid_type")]), + ], + ids=["array", "group", "array-with-a-problem", "group-with-a-problem"], +) +def test_a_node_is_read_as_the_node_its_node_type_says( + document: dict[str, object], + reading: type[ZarrV3ArrayMetadataReading | ZarrV3GroupMetadataReading], + problems: list[tuple[tuple[str | int, ...], str]], +) -> None: + """As that node's own read reads it, and its model only when it has no problem.""" + read = read_node_metadata_v3(document) + assert type(read) is reading + assert [(p.loc, p.kind) for p in read.problems] == problems + assert [(p.loc, p.kind) for p in validate_node_metadata_v3(document)] == problems + assert (read.metadata is not None) is (len(problems) == 0) + + +@pytest.mark.parametrize( + ("node_type", "kind"), + [("dataset", "invalid_value"), (5, "invalid_type"), (None, "invalid_type")], + ids=["another-kind", "a-number", "null"], +) +def test_error_a_node_type_the_spec_does_not_define(node_type: object, kind: str) -> None: + """Nothing else of the document is read but its `zarr_format`, which is 3 here, so nothing else is judged.""" + read = read_node_metadata_v3({**_array(), "node_type": node_type, "shape": "not a shape"}) + assert isinstance(read, ZarrV3UnknownNodeReading) + assert [(p.loc, p.kind) for p in read.problems] == [(("node_type",), kind)] + assert read.metadata is None + assert list(read.fields()) == [] + + +def test_error_a_document_without_a_node_type() -> None: + document = {key: value for key, value in _array().items() if key != "node_type"} + read = read_node_metadata_v3(document) + assert isinstance(read, ZarrV3UnknownNodeReading) + assert [(p.loc, p.kind) for p in read.problems] == [(("node_type",), "missing_key")] + + +@pytest.mark.parametrize( + ("document", "problems"), + [ + # zarr-python 2's draft of v3 (zarr-python#2982): a root `zarr.json` + # naming its format by URL, and no node type. + ( + { + "zarr_format": "https://purl.org/zarr/spec/protocol/core/3.0", + "metadata_encoding": "https://purl.org/zarr/spec/protocol/core/3.0", + "metadata_key_suffix": ".json", + "extensions": [], + }, + [(("zarr_format",), "invalid_type"), (("node_type",), "missing_key")], + ), + ( + {"zarr_format": 2, "shape": [4], "chunks": [4], "dtype": "|u1"}, + [(("zarr_format",), "invalid_value"), (("node_type",), "missing_key")], + ), + ( + {"node_type": "dataset"}, + [(("zarr_format",), "missing_key"), (("node_type",), "invalid_value")], + ), + ], + ids=["v3-draft", "v2", "no-format"], +) +def test_error_a_document_of_another_format_says_so( + document: dict[str, object], problems: list[tuple[tuple[str | int, ...], str]] +) -> None: + read = read_node_metadata_v3(document) + assert isinstance(read, ZarrV3UnknownNodeReading) + assert [(p.loc, p.kind) for p in read.problems] == problems + # The validator reads a node's type as the reader does. + assert [(p.loc, p.kind) for p in validate_node_metadata_v3(document)] == problems + + +def test_error_a_node_that_is_not_an_object() -> None: + read = read_node_metadata_v3([_array()]) + assert isinstance(read, ZarrV3UnknownNodeReading) + assert [(p.loc, p.kind) for p in read.problems] == [((), "invalid_type")] + + +@pytest.mark.parametrize( + "model", + [ + ZarrV3ArrayMetadata.create_default(shape=(4,)), + ZarrV3GroupMetadata.create_default(attributes={"a": 1}), + ], + ids=["array", "group"], +) +def test_a_node_is_built_as_the_model_its_node_type_says( + model: ZarrV3ArrayMetadata | ZarrV3GroupMetadata, +) -> None: + """From its JSON and from a store's bytes alike, as the model's own class builds it.""" + built = node_metadata_from_key_value_v3(model.to_key_value()) + assert type(built) is type(model) + assert built == model + assert node_metadata_from_json_v3(model.to_json()) == model + + +def test_error_a_node_built_from_a_store_without_its_document() -> None: + with pytest.raises(MetadataValidationError) as raised: + node_metadata_from_key_value_v3({}) + assert [(p.loc, p.kind) for p in raised.value.problems] == [(("zarr.json",), "missing_key")] + + +def test_error_a_node_built_from_bytes_that_are_not_json() -> None: + with pytest.raises(MetadataValidationError) as raised: + node_metadata_from_key_value_v3({"zarr.json": b"{"}) + assert [(p.loc, p.kind) for p in raised.value.problems] == [(("zarr.json",), "invalid_json")] + + +def test_error_a_node_built_from_a_document_of_no_node_type() -> None: + document = json.dumps({**_array(), "node_type": "dataset"}).encode() + with pytest.raises(MetadataValidationError) as raised: + node_metadata_from_key_value_v3({"zarr.json": document}) + assert [(p.loc, p.kind) for p in raised.value.problems] == [(("node_type",), "invalid_value")] + + +def test_error_a_node_built_from_a_document_with_a_problem() -> None: + """The problems of the node its `node_type` says it is.""" + with pytest.raises(MetadataValidationError) as raised: + node_metadata_from_json_v3(_array(fill_value=300)) + assert [(p.loc, p.kind) for p in raised.value.problems] == [(("fill_value",), "invalid_value")] + + # --- ZarrV2ConsolidatedMetadata -------------------------------------------- @@ -369,6 +1117,16 @@ def test_consolidated_v2_verbatim_roundtrip() -> None: assert model.to_json() == doc +@pytest.mark.parametrize(("value", "kind"), [("1", "invalid_type"), (2, "invalid_value")]) +def test_error_consolidated_v2_format_other_than_1(value: object, kind: str) -> None: + document = {"zarr_consolidated_format": value, "metadata": {}} + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ConsolidatedMetadata.from_json(document) + assert [(p.loc, p.kind) for p in raised.value.problems] == [ + (("zarr_consolidated_format",), kind) + ] + + def test_consolidated_v2_key_value_roundtrip() -> None: """from_key_value(to_key_value()) is the identity for .zmetadata documents.""" model = ZarrV2ConsolidatedMetadata.from_json( @@ -381,10 +1139,10 @@ def test_consolidated_v2_lists_become_tuples() -> None: """from_json converts JSON arrays inside entries to tuples.""" doc = { "zarr_consolidated_format": 1, - "metadata": {"a/.zarray": {"shape": [2, 3]}}, + "metadata": {"a/.zattrs": {"shape": [2, 3]}}, } model = ZarrV2ConsolidatedMetadata.from_json(doc) - assert model.metadata == {"a/.zarray": {"shape": (2, 3)}} + assert model.metadata == {"a/.zattrs": {"shape": (2, 3)}} def test_consolidated_v2_envelope_validation() -> None: @@ -395,7 +1153,7 @@ def test_consolidated_v2_envelope_validation() -> None: def test_consolidated_v2_not_a_mapping() -> None: """from_json rejects a non-mapping .zmetadata document.""" - with pytest.raises(MetadataValidationError, match="expected a mapping"): + with pytest.raises(MetadataValidationError, match="expected an object"): ZarrV2ConsolidatedMetadata.from_json([1]) @@ -495,7 +1253,7 @@ def test_group_v3_validator_agrees_with_from_json_on_consolidated() -> None: }, ) for doc in bad_docs: - assert validate_group_metadata_v3(doc) != [], doc + assert validate_group_metadata_v3(doc) != (), doc with pytest.raises(MetadataValidationError): ZarrV3GroupMetadata.from_json(doc) @@ -559,25 +1317,21 @@ def test_group_must_understand_fields_partition() -> None: https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L1571-L1573 """ model = ZarrV3GroupMetadata.create_default( - extra_fields={ - "waived": {"name": "w", "must_understand": False}, - "implicit": {"name": "i"}, - } + waived={"name": "w", "must_understand": False}, implicit={"name": "i"} ) assert set(model.must_understand_fields) == {"implicit"} -def test_group_v3_null_consolidated_metadata_repaired_to_absence() -> None: - """consolidated_metadata: null was written by a historical zarr-python bug. - Those stores must remain readable, but the bug spelling is not honored: - it is read as absence (UNSET) and never written back — the round-trip - deliberately repairs the document rather than preserving the bug.""" +def test_error_a_null_consolidated_metadata_is_a_value_the_document_wrote() -> None: + """zarr-python 3.0 and 3.1 wrote `consolidated_metadata: null`; the spec says an object, and the package models nothing else as right: a reader of those stores strips the key first.""" null_doc = {"zarr_format": 3, "node_type": "group", "consolidated_metadata": None} - assert validate_group_metadata_v3(null_doc) == () - model = ZarrV3GroupMetadata.from_json(null_doc) - assert model.consolidated_metadata is UNSET - assert "consolidated_metadata" not in model.to_json() - assert model == ZarrV3GroupMetadata.from_json({"zarr_format": 3, "node_type": "group"}) + assert [(p.loc, p.kind) for p in validate_group_metadata_v3(null_doc)] == [ + (("consolidated_metadata",), "invalid_type") + ] + with pytest.raises(MetadataValidationError): + ZarrV3GroupMetadata.from_json(null_doc) + del null_doc["consolidated_metadata"] + assert ZarrV3GroupMetadata.from_json(null_doc).consolidated_metadata is UNSET # --- to_json shares no mutable state with the model ------------------------ @@ -586,13 +1340,16 @@ def test_group_v3_null_consolidated_metadata_repaired_to_absence() -> None: pytest.param( ZarrV3GroupMetadata.create_default( attributes={"a": {"b": [1]}}, - consolidated_metadata=ZarrV3ConsolidatedMetadata( - metadata={ - "child": ZarrV3ArrayMetadata.create_default(attributes={"x": {"y": 1}}), - "grp": ZarrV3GroupMetadata.create_default(attributes={"x": {"y": 1}}), - } - ), - extra_fields={"ext": {"must_understand": False, "cfg": {"x": [1]}}}, + consolidated_metadata={ + **_INLINE_ENVELOPE, + "metadata": { + "child": ZarrV3ArrayMetadata.create_default( + attributes={"x": {"y": 1}} + ).to_json(), + "grp": ZarrV3GroupMetadata.create_default(attributes={"x": {"y": 1}}).to_json(), + }, + }, + ext={"must_understand": False, "cfg": {"x": [1]}}, ), id="v3-group", ), @@ -602,12 +1359,22 @@ def test_group_v3_null_consolidated_metadata_repaired_to_absence() -> None: ), pytest.param( ZarrV3ConsolidatedMetadata( - metadata={"child": ZarrV3ArrayMetadata.create_default(attributes={"x": {"y": 1}})} + { + **_INLINE_ENVELOPE, + "metadata": { + "child": ZarrV3ArrayMetadata.create_default( + attributes={"x": {"y": 1}} + ).to_json() + }, + } ), id="v3-consolidated", ), pytest.param( - ZarrV2ConsolidatedMetadata(metadata={"a/.zarray": {"nested": {"x": [1]}}}), + # A `.zattrs` entry: user data, whose containers can nest. + ZarrV2ConsolidatedMetadata( + {"zarr_consolidated_format": 1, "metadata": {"a/.zattrs": {"nested": {"x": [1]}}}} + ), id="v2-consolidated", ), ] @@ -641,3 +1408,57 @@ def test_from_json_shares_no_mutable_state_with_its_input( baseline = copy.deepcopy(read.to_json()) mutate_nested_containers(document) assert read.to_json() == baseline + + +def test_a_group_writes_no_empty_attributes() -> None: + model = ZarrV3GroupMetadata.create_default() + assert model.attributes == {} + assert "attributes" not in model.to_json() + assert json.loads(model.to_key_value()["zarr.json"]) == {"zarr_format": 3, "node_type": "group"} + + +def test_error_a_listed_group_s_own_listing_holds_the_documents_the_group_lists() -> None: + """One node, one document: a group listed in consolidated metadata may list a node the group lists too, but with the document the group holds at the joined key, as it is read; a document that reads otherwise is a problem at the nested entry, in the validator, the group model and the models given as entries alike.""" + same = _array(shape=(4,), attributes={"a": 1}) + other = _array(shape=(9,), attributes={"a": 1}) + agreed = _group( + consolidated_metadata=_inline( + b=_group(consolidated_metadata=_inline(c=same)), **{"b/c": same} + ) + ) + assert validate_group_metadata_v3(agreed) == () + contradicted = _group( + consolidated_metadata=_inline( + b=_group(consolidated_metadata=_inline(c=other)), **{"b/c": same} + ) + ) + found = validate_group_metadata_v3(contradicted) + at = ("consolidated_metadata", "metadata", "b", "consolidated_metadata", "metadata", "c") + assert [(p.loc, p.kind) for p in found] == [(at, "invalid_value")] + assert "/b/c" in found[0].message + with pytest.raises(MetadataValidationError): + ZarrV3GroupMetadata(contradicted) + inner = ZarrV3GroupMetadata(_group(consolidated_metadata=_inline(c=other))) + with pytest.raises(MetadataValidationError) as raised: + ZarrV3GroupMetadata( + _group(consolidated_metadata=_inline(b=inner, **{"b/c": ZarrV3ArrayMetadata(same)})) + ) + assert [(p.loc, p.kind) for p in raised.value.problems] == [(at, "invalid_value")] + + +def test_the_group_guard_asks_tuples_of_the_documents_the_group_lists() -> None: + """`is_group_metadata_v3` says no to a group whose consolidated metadata holds an array document written with lists, as `is_array_metadata_v3` says no to that document: a guard narrows to the TypedDict, whose arrays are tuples, at every level.""" + array = _array() + listed = dict(array) + listed["shape"] = list(array["shape"]) # pyright: ignore[reportArgumentType] + assert not is_array_metadata_v3(listed) + group = _group(consolidated_metadata=_inline(a=listed)) + assert validate_group_metadata_v3(group) == () + assert not is_group_metadata_v3(group) + assert is_group_metadata_v3(parse_group_metadata_v3(group)) + nested = _group( + consolidated_metadata=_inline( + g=_group(consolidated_metadata=_inline(a=listed)), **{"g/a": array} + ) + ) + assert not is_group_metadata_v3(nested) diff --git a/packages/zarr-metadata/tests/model/test_pair.py b/packages/zarr-metadata/tests/model/test_pair.py new file mode 100644 index 0000000000..d30409a5a1 --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_pair.py @@ -0,0 +1,740 @@ +"""A v3 model is a document and the scope it was read in: the pair, and what follows from it.""" + +from __future__ import annotations + +import copy +import dataclasses +import operator +import pickle +from collections.abc import Callable, Iterator, Mapping +from typing import Any, cast + +import pytest + +from zarr_metadata._json import refine_json +from zarr_metadata.model import ( + UNSET, + MetadataValidationError, + ZarrV3ArrayMetadata, + ZarrV3ConsolidatedMetadata, + ZarrV3GroupMetadata, + ZarrV3GroupMetadataReading, + read_array_metadata_v3, + read_group_metadata_v3, +) +from zarr_metadata.v3.codec.bytes import BYTES_CODEC +from zarr_metadata.v3.codec.crc32c import Empty +from zarr_metadata.v3.codec.zstd import ZSTD_CODEC +from zarr_metadata.v3.definition import ( + CORE, + CORE_AND_EXTENSIONS, + AcceptedField, + CodecDefinition, + Context, + Nested, + ScopeConflictError, + UnclaimedField, + ValidationProblem, +) + +ARRAY: dict[str, Any] = { + "zarr_format": 3, + "node_type": "array", + "shape": [4], + "data_type": "uint8", + "fill_value": 0, + "chunk_grid": {"name": "regular", "configuration": {"chunk_shape": [2]}}, + "chunk_key_encoding": {"name": "default"}, + "codecs": ["bytes"], +} +MY_BYTES = CodecDefinition(name="bytes", configuration=Empty, kind="array_bytes", size="static") +"""A private `bytes`: another meaning under the core name, which takes no configuration.""" +PRIVATE = CORE.extended_with(MY_BYTES) +SPELLED_OUT: dict[str, Any] = { + **ARRAY, + "data_type": {"name": "uint8"}, + "codecs": [{"name": "bytes", "configuration": {}}], + "chunk_key_encoding": {"name": "default", "configuration": {"separator": "/"}}, +} + + +@pytest.mark.parametrize( + ("document", "context", "expected_context"), + [ + (ARRAY, None, CORE_AND_EXTENSIONS), + (ARRAY, CORE, CORE), + (SPELLED_OUT, Context.of(), Context.of()), + ], + ids=["default-scope", "core", "empty-scope"], +) +def test_a_model_is_its_document_read_in_its_scope( + document: dict[str, Any], context: Context | None, expected_context: Context +) -> None: + """A model built from a document holds that document, as written and refined, and the scope it was read in, `CORE_AND_EXTENSIONS` when none is given.""" + model = ZarrV3ArrayMetadata(document, context=context) + assert model.context == expected_context + assert model.to_json() == refine_json(document)[0] + + +@pytest.mark.parametrize( + ("document", "key", "written"), + [ + (ARRAY, "data_type", "uint8"), + (SPELLED_OUT, "data_type", {"name": "uint8"}), + (ARRAY, "codecs", ("bytes",)), + (SPELLED_OUT, "codecs", ({"name": "bytes", "configuration": {}},)), + ], + ids=["bare-data-type", "object-data-type", "bare-codec", "object-codec"], +) +def test_to_json_keeps_the_spelling_the_document_was_written_in( + document: dict[str, Any], key: str, written: object +) -> None: + """`to_json` writes each field as the document wrote it, not as a reader would respell it: the model is the document.""" + assert ZarrV3ArrayMetadata(document).to_json()[key] == written + + +def test_properties_are_what_the_reading_holds() -> None: + """The typed members -- fields as the scope read them, shape, fill value, attributes -- are views of the reading, read-only.""" + model = ZarrV3ArrayMetadata({**ARRAY, "attributes": {"a": [1]}, "acme": 1}) + assert isinstance(model.data_type, AcceptedField) + assert isinstance(model.codecs[0], AcceptedField) + assert model.shape == (4,) + assert model.fill_value == 0 + assert model.dimension_names is UNSET + assert isinstance(model.attributes, Mapping) + assert model.attributes == {"a": (1,)} + assert model.extra_fields == {"acme": 1} + assert (model.zarr_format, model.node_type) == (3, "array") + with pytest.raises(TypeError): + model.attributes["b"] = 1 # pyright: ignore[reportIndexIssue] + with pytest.raises(AttributeError): + model.shape = (5,) # pyright: ignore[reportAttributeAccessIssue] + assert isinstance(ZarrV3ArrayMetadata(ARRAY, context=Context.of()).data_type, UnclaimedField) + + +def test_a_reading_without_problems_builds_the_model_without_reading_again() -> None: + """`read_array_metadata_v3` hands its reading to the model it builds, so the model's `reading` is that reading, not a second one.""" + reading = read_array_metadata_v3(ARRAY) + assert reading.metadata is not None + assert reading.metadata.reading.pipeline is reading.pipeline + + +def test_error_a_document_with_a_problem_is_refused_at_construction() -> None: + """The constructor raises `MetadataValidationError` with every problem, as `from_json` does.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV3ArrayMetadata({**ARRAY, "fill_value": 300, "shape": [-1]}) + assert sorted(problem.loc for problem in raised.value.problems) == [ + ("fill_value",), + ("shape",), + ] + + +@pytest.mark.parametrize( + ("left", "right", "equal"), + [ + (ZarrV3ArrayMetadata(ARRAY), ZarrV3ArrayMetadata(SPELLED_OUT), True), + (ZarrV3ArrayMetadata(ARRAY, context=CORE), ZarrV3ArrayMetadata(ARRAY), True), + (ZarrV3ArrayMetadata(ARRAY), ZarrV3ArrayMetadata(ARRAY, context=PRIVATE), False), + (ZarrV3ArrayMetadata(ARRAY), ZarrV3ArrayMetadata(ARRAY, context=Context.of()), False), + (ZarrV3ArrayMetadata(ARRAY), ZarrV3ArrayMetadata({**ARRAY, "attributes": {}}), True), + ], + ids=["spellings", "unused-definitions", "private-bytes", "unclaimed", "empty-attributes"], +) +def test_models_are_equal_by_what_their_documents_mean( + left: ZarrV3ArrayMetadata, right: ZarrV3ArrayMetadata, equal: bool +) -> None: + """Two models are equal when every field reads the same by the same definition and the rest is the same JSON: spelling and unused definitions do not matter, a private definition under a core name does, and so does a name read against one left unclaimed.""" + assert (left == right) is equal + if equal: + assert hash(left) == hash(right) + + +@pytest.mark.parametrize( + "context", [None, CORE, PRIVATE, Context.of()], ids=["default", "core", "private", "empty"] +) +def test_a_model_round_trips_through_its_document_in_its_scope(context: Context | None) -> None: + """`from_json(m.to_json(), context=m.context) == m`: the document and the scope determine the model, and pickle carries both.""" + model = ZarrV3ArrayMetadata(ARRAY, context=context) + assert ZarrV3ArrayMetadata(model.to_json(), context=model.context) == model + loaded = pickle.loads(pickle.dumps(model)) + assert loaded == model + assert loaded.context == model.context + assert loaded.to_json() == model.to_json() + + +def test_error_a_model_whose_scope_does_not_pickle_says_so() -> None: + """A scope holding a definition with a local function does not pickle, and the model raises the error pickling a `Context` raises.""" + local = CodecDefinition( + name="acme.c", + configuration=Empty, + kind="bytes_bytes", + size="dynamic", + rules=lambda configuration, nested: iter(()), + ) + model = ZarrV3ArrayMetadata( + {**ARRAY, "codecs": ["bytes", "acme.c"]}, context=CORE.extended_with(local) + ) + with pytest.raises((pickle.PicklingError, AttributeError)): + pickle.dumps(model) + + +FLOAT: dict[str, Any] = { + **ARRAY, + "data_type": "float32", + "fill_value": "0x7fc00000", + "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}], +} +SHARDED: dict[str, Any] = { + **ARRAY, + "codecs": [ + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [1], + "codecs": ["bytes"], + "index_codecs": [ + {"name": "bytes", "configuration": {"endian": "little"}}, + "crc32c", + ], + }, + } + ], +} + + +@pytest.mark.parametrize( + ("model", "members", "expected"), + [ + ( + ZarrV3ArrayMetadata(ARRAY), + {"attributes": {"a": 1}}, + ZarrV3ArrayMetadata({**ARRAY, "attributes": {"a": 1}}), + ), + ( + ZarrV3ArrayMetadata({**ARRAY, "attributes": {"a": 1}}), + {"attributes": UNSET}, + ZarrV3ArrayMetadata(ARRAY), + ), + ( + ZarrV3ArrayMetadata(ARRAY, context=PRIVATE), + {"attributes": {"a": 1}}, + ZarrV3ArrayMetadata({**ARRAY, "attributes": {"a": 1}}, context=PRIVATE), + ), + ( + ZarrV3ArrayMetadata(ARRAY), + { + "shape": (6,), + "chunk_grid": {"name": "regular", "configuration": {"chunk_shape": (3,)}}, + }, + ZarrV3ArrayMetadata( + { + **ARRAY, + "shape": [6], + "chunk_grid": {"name": "regular", "configuration": {"chunk_shape": [3]}}, + } + ), + ), + ], + ids=["set", "unset", "update-keeps-scope", "members-together"], +) +def test_update_reads_new_members_in_the_models_own_scope( + model: ZarrV3ArrayMetadata, members: dict[str, Any], expected: ZarrV3ArrayMetadata +) -> None: + """`update` puts JSON members in place of the document's, `UNSET` removing one, and reads the result in the model's own scope, so a model read privately stays private.""" + updated = model.update(**members) + assert updated == expected + assert updated.context == model.context + + +def test_error_update_refuses_a_document_with_a_problem() -> None: + """`update` raises `MetadataValidationError` when the members make an invalid document, as the constructor does: two dimension names for one dimension.""" + with pytest.raises(MetadataValidationError): + ZarrV3ArrayMetadata(ARRAY).update(dimension_names=("x", "y")) + + +@pytest.mark.parametrize( + ("model", "context", "gained"), + [ + (ZarrV3ArrayMetadata(ARRAY, context=Context.of()), CORE, True), + (ZarrV3ArrayMetadata(ARRAY, context=CORE), CORE_AND_EXTENSIONS, False), + (ZarrV3ArrayMetadata(FLOAT, context=Context.of()), CORE, True), + (ZarrV3ArrayMetadata(SHARDED, context=Context.of()), CORE, True), + ], + ids=["gain", "nothing-to-gain", "gain-data-type-with-fill-value", "gain-of-a-shard"], +) +def test_refined_in_moves_a_model_up_the_order( + model: ZarrV3ArrayMetadata, context: Context, gained: bool +) -> None: + """`refined_in` reads the document in a scope that claims what this one left unclaimed and contradicts nothing; the result refines the model, keeps its document, and is the model itself when the scope reads nothing otherwise.""" + refined = model.refined_in(context) + assert refined.context == context + assert refined.to_json() == model.to_json() + assert refined.refines(model) + assert (refined == model) is not gained + if not gained: + # Not read again: the reading's pipeline is the one read. + assert refined.reading.pipeline is model.reading.pipeline + + +@pytest.mark.parametrize( + ("model", "context"), + [ + (ZarrV3ArrayMetadata(ARRAY), PRIVATE), + (ZarrV3ArrayMetadata(ARRAY, context=CORE), Context.of()), + ], + ids=["conflict", "loss"], +) +def test_error_refined_in_refuses_a_conflict_or_a_loss( + model: ZarrV3ArrayMetadata, context: Context +) -> None: + """`refined_in` raises `ScopeConflictError` naming each name the scope reads by another definition, or by none.""" + with pytest.raises(ScopeConflictError) as raised: + model.refined_in(context) + assert "bytes" in [conflict.key[1] for conflict in raised.value.conflicts] + + +def test_with_context_reads_the_document_in_any_scope() -> None: + """`with_context` reads the same document in another scope, whatever that changes: a private `bytes` is read as such, and a scope that claims nothing leaves every name unclaimed.""" + model = ZarrV3ArrayMetadata(ARRAY) + private = model.with_context(PRIVATE) + assert private.context == PRIVATE + assert private.codecs[0].definition == MY_BYTES + assert isinstance(model.with_context(Context.of()).codecs[0], UnclaimedField) + assert model.with_context(None).context == CORE_AND_EXTENSIONS + + +def test_error_with_context_refuses_a_document_the_scope_reads_with_a_problem() -> None: + """`with_context` raises `MetadataValidationError` when the document has a problem in the new scope: a gzip `level` the core definition refuses.""" + loose = ZarrV3ArrayMetadata( + {**ARRAY, "codecs": ["bytes", {"name": "gzip", "configuration": {"level": 12}}]}, + context=Context.of(), + ) + with pytest.raises(MetadataValidationError): + loose.with_context(CORE) + + +@pytest.mark.parametrize( + ("upper", "lower", "expected"), + [ + (ZarrV3ArrayMetadata(ARRAY), ZarrV3ArrayMetadata(ARRAY, context=Context.of()), True), + (ZarrV3ArrayMetadata(ARRAY, context=Context.of()), ZarrV3ArrayMetadata(ARRAY), False), + (ZarrV3ArrayMetadata(ARRAY), ZarrV3ArrayMetadata(SPELLED_OUT), True), + (ZarrV3ArrayMetadata(ARRAY, context=PRIVATE), ZarrV3ArrayMetadata(ARRAY), False), + (ZarrV3ArrayMetadata({**ARRAY, "attributes": {"a": 1}}), ZarrV3ArrayMetadata(ARRAY), False), + ( + ZarrV3ArrayMetadata(FLOAT, context=CORE), + ZarrV3ArrayMetadata(FLOAT, context=Context.of()), + True, + ), + ( + ZarrV3ArrayMetadata(SHARDED, context=CORE), + ZarrV3ArrayMetadata(SHARDED, context=Context.of()), + True, + ), + ( + ZarrV3ArrayMetadata(FLOAT, context=CORE), + ZarrV3ArrayMetadata({**FLOAT, "fill_value": "banana"}, context=Context.of()), + False, + ), + ], + ids=[ + "gain", + "loss", + "equal", + "conflict", + "other-members-differ", + "fill-value-spelled-by-the-informed-side", + "gain-of-a-field-holding-fields", + "fill-value-the-gained-definition-refuses", + ], +) +def test_refines_orders_models_by_information( + upper: ZarrV3ArrayMetadata, lower: ZarrV3ArrayMetadata, expected: bool +) -> None: + """A model refines another when every field refines its counterpart -- the fields a field holds too, so a shard gained is a gain -- and every other member is the same, the fill value compared as the more informed data type spells it; a fill value that definition refuses is no refinement, and no error.""" + assert upper.refines(lower) is expected + + +GROUP: dict[str, Any] = {"zarr_format": 3, "node_type": "group", "attributes": {"g": 1}} +CONSOLIDATED: dict[str, Any] = { + **GROUP, + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": { + "a": ARRAY, + "b": { + **GROUP, + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": {"c": SPELLED_OUT}, + }, + }, + "b/c": SPELLED_OUT, + }, + }, +} + + +def test_a_group_is_its_document_and_scope_and_its_nested_models_share_them() -> None: + """A group model is the pair, and each document its consolidated metadata holds is a model of the same scope, built from the group's one read, writing its document as written.""" + group = ZarrV3GroupMetadata(CONSOLIDATED, context=CORE) + held = group.consolidated_metadata + assert isinstance(held, ZarrV3ConsolidatedMetadata) + assert set(held.metadata) == {"a", "b", "b/c"} + for node in held.metadata.values(): + assert node.context is group.context + assert held.metadata["b/c"].to_json()["data_type"] == {"name": "uint8"} + assert group.to_json() == refine_json(CONSOLIDATED)[0] + assert ZarrV3GroupMetadata(GROUP).consolidated_metadata is UNSET + + +def test_a_group_reads_each_nested_field_once() -> None: + """Building a group with consolidated metadata asks a definition's rules once per nested field: the models are built from the read, not read again.""" + calls: list[int] = [] + + def counted(configuration: object, nested: Nested) -> Iterator[ValidationProblem]: + calls.append(1) + return iter(()) + + scope = CORE.extended_with(dataclasses.replace(BYTES_CODEC, rules=counted)) + ZarrV3GroupMetadata(CONSOLIDATED, context=scope) + assert len(calls) == 3 # `a`, `b/c`, and `c` in `b`'s own listing, each one bytes codec + + +@pytest.mark.parametrize( + ("left", "right", "equal"), + [ + (ZarrV3GroupMetadata(CONSOLIDATED), ZarrV3GroupMetadata(CONSOLIDATED, context=CORE), True), + ( + ZarrV3GroupMetadata(CONSOLIDATED), + ZarrV3GroupMetadata(CONSOLIDATED, context=PRIVATE), + False, + ), + (ZarrV3GroupMetadata(GROUP), ZarrV3GroupMetadata({**GROUP, "attributes": {"g": 2}}), False), + ], + ids=["unused-definitions", "private-bytes-inside", "attributes"], +) +def test_groups_are_equal_by_what_their_documents_mean( + left: ZarrV3GroupMetadata, right: ZarrV3GroupMetadata, equal: bool +) -> None: + """A group compares by its attributes, extra fields and each nested model's meaning; equal groups hash alike.""" + assert (left == right) is equal + if equal: + assert hash(left) == hash(right) + + +@pytest.mark.parametrize("document", [GROUP, CONSOLIDATED], ids=["group", "consolidated"]) +def test_a_group_round_trips_through_its_document_in_its_scope(document: dict[str, Any]) -> None: + """`from_json(g.to_json(), context=g.context) == g`, and pickle carries the pair.""" + group = ZarrV3GroupMetadata(document, context=CORE) + assert ZarrV3GroupMetadata(group.to_json(), context=group.context) == group + assert pickle.loads(pickle.dumps(group)) == group + + +def test_group_update_with_context_and_refined_in_behave_as_the_arrays_do() -> None: + """`update` reads in the group's scope and keeps its consolidated metadata unless given; `refined_in` moves every nested model up the order; `with_context` reads all of it in another scope.""" + group = ZarrV3GroupMetadata(CONSOLIDATED, context=Context.of()) + assert group.update(attributes={"g": 2}).consolidated_metadata == group.consolidated_metadata + refined = group.refined_in(CORE) + assert refined.refines(group) + assert refined.context == CORE + held = refined.consolidated_metadata + assert isinstance(held, ZarrV3ConsolidatedMetadata) + array = held.metadata["a"] + assert isinstance(array, ZarrV3ArrayMetadata) + assert isinstance(array.codecs[0], AcceptedField) + with pytest.raises(ScopeConflictError): + ZarrV3GroupMetadata(CONSOLIDATED).refined_in(PRIVATE) + private = group.with_context(PRIVATE).consolidated_metadata + assert isinstance(private, ZarrV3ConsolidatedMetadata) + private_array = private.metadata["a"] + assert isinstance(private_array, ZarrV3ArrayMetadata) + assert private_array.codecs[0].definition == MY_BYTES + + +def test_consolidated_metadata_reads_on_its_own() -> None: + """`ZarrV3ConsolidatedMetadata(member, context)` reads the member as a group's read reads it, each document a model of that scope.""" + member = CONSOLIDATED["consolidated_metadata"] + held = ZarrV3ConsolidatedMetadata(member, context=CORE) + assert held == ZarrV3GroupMetadata(CONSOLIDATED, context=CORE).consolidated_metadata + assert held.to_json() == refine_json(member)[0] + assert ZarrV3ConsolidatedMetadata.from_json(member) == ZarrV3ConsolidatedMetadata(member) + + +def test_error_a_group_with_a_nested_problem_is_refused_at_the_nested_path() -> None: + """A nested document's problem is the group's, located under `consolidated_metadata.metadata.`.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV3GroupMetadata( + { + **CONSOLIDATED, + "consolidated_metadata": { + **CONSOLIDATED["consolidated_metadata"], + "metadata": {"a": {**ARRAY, "shape": [-1]}}, + }, + } + ) + assert raised.value.problems[0].loc == ("consolidated_metadata", "metadata", "a", "shape") + + +def test_the_held_field_machinery_is_gone() -> None: + """A v3 model is built only by reading its document, so nothing in the package takes fields read already, or re-reads a model in an empty scope: `NO_SCOPE`, `overlapping` and `held` are gone.""" + import inspect + + import zarr_metadata.model._validation as validation + + assert not hasattr(validation, "NO_SCOPE") + assert not hasattr(validation, "overlapping") + assert "held" not in inspect.signature(validation.read_array_v3).parameters + + +def test_error_refined_in_refuses_a_gain_that_surfaces_a_problem() -> None: + """A scope that claims a name this one left unclaimed may refuse what was written under it: `refined_in` raises `MetadataValidationError`, the document having a problem in that scope.""" + loose = ZarrV3ArrayMetadata({**ARRAY, "fill_value": "banana"}, context=Context.of()) + with pytest.raises(MetadataValidationError) as raised: + loose.refined_in(CORE) + assert [problem.loc for problem in raised.value.problems] == [("fill_value",)] + + +def test_a_scope_conflict_says_where_each_conflict_sits() -> None: + """`refined_in` names each conflict with where the field sits in the document, as a problem is located: a loss of `bytes` at `codecs.0`, in the shard too.""" + model = ZarrV3ArrayMetadata(SHARDED, context=CORE) + with pytest.raises(ScopeConflictError) as raised: + model.refined_in(Context.of()) + assert sorted((conflict.key[1], conflict.loc) for conflict in raised.value.conflicts) == [ + ("bytes", ("codecs", 0, "configuration", "codecs", 0)), + ("bytes", ("codecs", 0, "configuration", "index_codecs", 0)), + ("crc32c", ("codecs", 0, "configuration", "index_codecs", 1)), + ("default", ("chunk_key_encoding",)), + ("regular", ("chunk_grid",)), + ("sharding_indexed", ("codecs", 0)), + ("uint8", ("data_type",)), + ] + + +@pytest.mark.parametrize( + "build", + [ + lambda: ZarrV3GroupMetadata(CONSOLIDATED, context=CORE_AND_EXTENSIONS), + lambda: read_group_metadata_v3(CONSOLIDATED, context=CORE_AND_EXTENSIONS).metadata, + ], + ids=["constructor", "reader"], +) +def test_a_models_reading_holds_the_model_and_with_context_moves_the_whole_tree( + build: Callable[[], ZarrV3GroupMetadata | None], +) -> None: + """However a group was built, its reading holds it, each nested reading holds the nested model, and `with_context` into a scope that reads every claim identically moves every nested model, the nested ones' too, to the new scope without reading again.""" + group = build() + assert group is not None + assert group.reading.metadata is group + nested = group.consolidated_metadata + assert isinstance(nested, ZarrV3ConsolidatedMetadata) + for path, node in nested.metadata.items(): + assert group.reading.consolidated[path].metadata is node + assert node.reading.metadata is node + scope = CORE.extended_with(ZSTD_CODEC) + moved = group.with_context(scope) + assert moved.reading is not group.reading + held = moved.consolidated_metadata + assert isinstance(held, ZarrV3ConsolidatedMetadata) + assert held.context is scope + before = nested.metadata["a"] + after = held.metadata["a"] + assert isinstance(before, ZarrV3ArrayMetadata) + assert isinstance(after, ZarrV3ArrayMetadata) + assert after.reading.pipeline is before.reading.pipeline # not read again + for node in held.metadata.values(): + assert node.context is scope + assert node.reading.metadata is node + inner = held.metadata["b"] + assert isinstance(inner, ZarrV3GroupMetadata) + innermost = inner.consolidated_metadata + assert isinstance(innermost, ZarrV3ConsolidatedMetadata) + assert innermost.context is scope + assert innermost.metadata["c"].context is scope + assert innermost.metadata["c"].reading.metadata is innermost.metadata["c"] + + +def test_consolidated_metadata_refines_nothing_of_another_type() -> None: + """`refines` of consolidated metadata says False of what is not consolidated metadata, as the array's and group's do, rather than raising.""" + held = ZarrV3GroupMetadata(CONSOLIDATED).consolidated_metadata + assert isinstance(held, ZarrV3ConsolidatedMetadata) + assert held.refines(cast("Any", ZarrV3GroupMetadata(GROUP))) is False + + +@pytest.mark.parametrize( + ("document", "change"), + [ + ( + {**ARRAY, "attributes": {"a": {"b": 1}}}, + lambda model: operator.setitem(model.attributes["a"], "b", 2), + ), + ( + {**ARRAY, "acme": {"x": 1, "must_understand": False}}, + lambda model: operator.setitem(model.extra_fields["acme"], "x", 2), + ), + ( + { + **ARRAY, + "data_type": { + "name": "struct", + "configuration": {"fields": [{"name": "a", "data_type": "uint8"}]}, + }, + "fill_value": {"a": 0}, + }, + lambda model: operator.setitem(model.fill_value, "a", 7), + ), + ], + ids=["attributes", "extra-fields", "fill-value"], +) +def test_a_models_views_cannot_be_changed_in_place( + document: dict[str, Any], change: Callable[[ZarrV3ArrayMetadata], None] +) -> None: + """What a model shows -- attributes, extra fields, a fill value -- is read-only at every level, so a model cannot be put in a state its document, its key and `refines` disagree about.""" + model = ZarrV3ArrayMetadata(document) + same = ZarrV3ArrayMetadata(document) + with pytest.raises(TypeError): + change(model) + assert model == same + assert model.to_json() == same.to_json() + assert model.refines(same) + + +@pytest.mark.parametrize( + "problem", + [ + {"attributes": {1: "x"}}, + {"attributes": 5}, + { + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": {"a": ARRAY, "b": {**GROUP, "attributes": {"s": {1, 2}}}}, + } + }, + { + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": {"a": ARRAY, "b": {**ARRAY, "shape": "x"}}, + } + }, + ], + ids=["non-string-key", "attributes-not-an-object", "sibling-not-json", "sibling-invalid"], +) +def test_a_reading_holds_a_model_of_each_nested_document_without_a_problem( + problem: dict[str, Any], +) -> None: + """A group document with a problem still holds, in its reading, a model of each document its consolidated metadata holds that has no problem, whatever the problem elsewhere is: one a reader walks past, or one it refuses.""" + member = {"kind": "inline", "must_understand": False, "metadata": {"a": ARRAY}} + document = {**GROUP, "consolidated_metadata": member, **problem} + reading = read_group_metadata_v3(document) + assert len(reading.problems) != 0 + assert reading.metadata is None + held = reading.consolidated["a"].metadata + assert isinstance(held, ZarrV3ArrayMetadata) + assert held == ZarrV3ArrayMetadata(ARRAY) + + +def test_a_scope_conflict_inside_consolidated_metadata_is_located_there() -> None: + """A group's `refined_in` locates a conflict in a document its consolidated metadata holds under that document's path, in a listing a listed group holds too.""" + group = ZarrV3GroupMetadata(CONSOLIDATED, context=CORE) + with pytest.raises(ScopeConflictError) as raised: + group.refined_in(Context.of()) + locs = {conflict.loc for conflict in raised.value.conflicts} + assert ("consolidated_metadata", "metadata", "a", "codecs", 0) in locs + assert ( + "consolidated_metadata", + "metadata", + "b", + "consolidated_metadata", + "metadata", + "c", + "codecs", + 0, + ) in locs + + +def test_a_models_reading_cannot_be_changed_in_place() -> None: + """What a model's reading holds of the documents its consolidated metadata holds is read-only, so nothing planted there is taken up by `with_context`.""" + group = ZarrV3GroupMetadata(CONSOLIDATED) + planted = ZarrV3ArrayMetadata({**ARRAY, "shape": [9]}).reading + with pytest.raises(TypeError): + operator.setitem(cast("Any", group.reading.consolidated), "a", planted) + moved = group.with_context(CORE_AND_EXTENSIONS) + assert moved == ZarrV3GroupMetadata(CONSOLIDATED) + + +@pytest.mark.parametrize( + "fault", + [{"attributes": "bad"}, {"attributes": {1: "bad"}}, {"attributes": {"s": {1, 2}}}], + ids=["refused", "non-string-key", "not-json"], +) +def test_a_reading_holds_a_model_of_each_healthy_document_in_a_listed_groups_own_listing( + fault: dict[str, Any], +) -> None: + """A listed group with a problem of its own -- one the reader refuses, or one it walks past -- still holds, in its reading, a model of each document in its own listing that has no problem, as the top group does.""" + inline = {"kind": "inline", "must_understand": False} + listed = { + **GROUP, + **fault, + "consolidated_metadata": { + **inline, + "metadata": {"x": ARRAY, "y": {**ARRAY, "shape": [-1]}}, + }, + } + document = { + **GROUP, + "consolidated_metadata": { + **inline, + "metadata": {"a": listed, "a/x": ARRAY, "a/y": {**ARRAY, "shape": [-1]}}, + }, + } + reading = read_group_metadata_v3(document) + assert reading.metadata is None + nested = reading.consolidated["a"] + assert isinstance(nested, ZarrV3GroupMetadataReading) + assert nested.metadata is None + held = nested.consolidated["x"].metadata + assert isinstance(held, ZarrV3ArrayMetadata) + assert held == ZarrV3ArrayMetadata(ARRAY) + + +@pytest.mark.parametrize("document", [GROUP, CONSOLIDATED], ids=["group", "consolidated"]) +def test_a_group_reading_pickles_and_copies(document: dict[str, Any]) -> None: + """A group's reading pickles and deep-copies, with the documents its consolidated metadata holds and the model it built, and compares equal afterwards, as an array's does.""" + reading = read_group_metadata_v3(document) + for again in (pickle.loads(pickle.dumps(reading)), copy.deepcopy(reading)): + assert again == reading + assert again.metadata == reading.metadata + assert set(again.consolidated) == set(reading.consolidated) + # One model per document still: the reading comes back through + # the model it holds, which reads once. + assert again.metadata is not None + assert again.metadata.reading is again + for path, nested in again.consolidated.items(): + held = again.metadata.consolidated_metadata + assert isinstance(held, ZarrV3ConsolidatedMetadata) + assert nested.metadata is held.metadata[path] + + +def test_an_array_reading_pickles_through_its_model() -> None: + """An array's reading that holds a model pickles and deep-copies as that model does, and comes back as the model's own reading.""" + reading = read_array_metadata_v3(ARRAY) + for again in (pickle.loads(pickle.dumps(reading)), copy.deepcopy(reading)): + assert again == reading + assert again.metadata is not None + assert again.metadata.reading is again + + +def test_error_a_models_fields_cannot_be_changed_in_place() -> None: + """The fields a model hands out -- a codec's configuration, the fields a shard holds -- are read-only, so `codecs ==` and `refines` cannot drift from `==`.""" + model = ZarrV3ArrayMetadata(ARRAY) + same = ZarrV3ArrayMetadata(ARRAY) + codec = model.codecs[0] + assert isinstance(codec, AcceptedField) + with pytest.raises(TypeError): + codec.configuration["endian"] = "big" # pyright: ignore[reportIndexIssue] + assert model.codecs == same.codecs + assert model.refines(same) diff --git a/packages/zarr-metadata/tests/model/test_pair_v2.py b/packages/zarr-metadata/tests/model/test_pair_v2.py new file mode 100644 index 0000000000..b09502401b --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_pair_v2.py @@ -0,0 +1,472 @@ +"""A v2 model is its document and the scope it was read in, as the v3 models are.""" + +from __future__ import annotations + +import copy +import dataclasses +import pickle +from typing import Any, cast + +import pytest + +from zarr_metadata._sentinel import UNSET +from zarr_metadata.model import ( + MetadataValidationError, + ZarrV2ArrayMetadata, + ZarrV2ConsolidatedMetadata, + ZarrV2GroupMetadata, + read_array_metadata_v2, +) +from zarr_metadata.v2.codec.compression import ZLIB_V2 +from zarr_metadata.v2.data_type.scalar import FLOAT_V2, UINT_V2 +from zarr_metadata.v2.definition import ( + CORE_V2, + AcceptedField, + Context, + ScopeConflictError, + UnclaimedField, + ZarrV2CodecDefinition, + ZarrV2DataTypeDefinition, +) +from zarr_metadata.v3.definition import ( + EmptyConfiguration, +) + +Loc = tuple[str | int, ...] +ARRAY: dict[str, Any] = { + "zarr_format": 2, + "shape": [4], + "chunks": [2], + "dtype": " dict[str, Any]: + """`ARRAY` with `changes`, a member given as `UNSET` left out.""" + return {key: value for key, value in {**ARRAY, **changes}.items() if value is not UNSET} + + +def test_a_model_is_its_document_read_in_its_scope() -> None: + """The constructor reads the document in the scope given, `CORE_V2` by default; `to_json` is the document refined (arrays as tuples, `dimension_separator` put in when missing), sharing nothing with the model; `context` is the scope; the reading holds the model.""" + model = ZarrV2ArrayMetadata(ARRAY) + assert model.context is CORE_V2 + document = model.to_json() + assert document["shape"] == (4,) + assert document.get("attributes") == {"a": (1, 2)} + without = {k: v for k, v in ARRAY.items() if k != "dimension_separator"} + assert ZarrV2ArrayMetadata(without).to_json().get("dimension_separator") == "." + assert model.reading.metadata is model + assert read_array_metadata_v2(ARRAY).metadata == model + assert ZarrV2ArrayMetadata(ARRAY, SMALL).context is SMALL + + +def test_properties_are_what_the_reading_holds() -> None: + """`dtype`, `compressor` and `filters` are the fields as the scope read them, None where `null` is written; `shape`, `chunks`, `fill_value`, `order`, `dimension_separator`, `attributes` and `extra_fields` are the members as the read refined them, read-only; `claims` is keyed as the scope files them.""" + model = ZarrV2ArrayMetadata({**ARRAY, "filters": [{"id": "x"}], "extra": [1]}) + assert isinstance(model.dtype, AcceptedField) + assert model.dtype.definition is FLOAT_V2 + assert isinstance(model.compressor, AcceptedField) + assert model.compressor.definition is ZLIB_V2 + assert model.filters is not None + assert isinstance(model.filters[0], UnclaimedField) + members = (model.shape, model.chunks, model.fill_value, model.order, model.dimension_separator) + assert members == ((4,), (2,), 0, "C", ".") + assert model.attributes == {"a": (1, 2)} + assert model.extra_fields == {"extra": (1,)} + assert (ZarrV2DataTypeDefinition, "float") in model.claims + with pytest.raises(TypeError): + cast("dict[str, object]", model.attributes)["b"] = 1 + assert ZarrV2ArrayMetadata({**ARRAY, "compressor": None}).compressor is None + assert ZarrV2ArrayMetadata(_document({"attributes": UNSET})).attributes is UNSET + + +@pytest.mark.parametrize( + ("left", "right", "same"), + [ + ({}, {"shape": (4,), "attributes": {"a": (1, 2)}}, True), + ({"dtype": "f4"}, False), + ({"attributes": {"a": [1, 2]}}, {"attributes": {"a": [2, 1]}}, False), + ({"attributes": UNSET}, {"attributes": {}}, False), + ({"extra": 1}, {}, False), + ], + ids=[ + "same", + "int-for-float", + "nan", + "spelling", + "key-order", + "signed-zero", + "byte-order", + "attributes", + "zattrs-presence", + "extra", + ], +) +def test_models_are_equal_by_what_their_documents_mean( + left: dict[str, Any], right: dict[str, Any], same: bool +) -> None: + """Two models are one array when their documents mean the same in their scopes: a typestr spelled two ways or a fill value an integer or a float is one, a byte order or a `.zattrs` present or not is two; equal models hash alike.""" + one, other = ZarrV2ArrayMetadata(_document(left)), ZarrV2ArrayMetadata(_document(right)) + assert (one == other) is same + if same: + assert hash(one) == hash(other) + + +def test_a_model_read_in_two_scopes_that_read_it_alike_is_one_model() -> None: + """Equality is by interpretation: the same document read by a private `zlib` with no parameters and by the core one are two models, and read in two scopes that file the same `zlib` are one.""" + document = {**ARRAY, "compressor": {"id": "zlib"}} + assert ZarrV2ArrayMetadata(document) != ZarrV2ArrayMetadata(document, PRIVATE) + same = Context.of(*CORE_V2.definitions()) + assert ZarrV2ArrayMetadata(document) == ZarrV2ArrayMetadata(document, same) + + +@pytest.mark.parametrize("context", [None, SMALL, PRIVATE], ids=["core", "small", "private"]) +def test_a_model_round_trips_through_its_document_pickle_and_copy(context: Context | None) -> None: + """`from_json(m.to_json(), context=m.context)`, `from_key_value(m.to_key_value())`, `pickle` and `copy.deepcopy` give an equal model in the same scope, and a pickled reading comes back as its model's own reading.""" + model = ZarrV2ArrayMetadata(ARRAY, context) + assert ZarrV2ArrayMetadata.from_json(model.to_json(), context=model.context) == model + restored = ZarrV2ArrayMetadata.from_key_value(model.to_key_value(), context=model.context) + assert restored == model + loaded = pickle.loads(pickle.dumps(model)) + assert loaded == model + assert loaded.context == model.context + assert copy.deepcopy(model) == model + assert pickle.loads(pickle.dumps(model.reading)).metadata == model + + +@pytest.mark.parametrize( + ("changes", "expect"), + [ + ({"shape": [8], "chunks": [4]}, {"shape": (8,), "chunks": (4,)}), + ({"dtype": " None: + """`update` puts JSON members in place of the document's, leaves out one given as `UNSET`, and reads the result in the model's own scope; the model is unchanged.""" + model = ZarrV2ArrayMetadata({**ARRAY, "extra": 0}, PRIVATE) + changed = model.update(**changes) + assert changed.context is PRIVATE + document = changed.to_json() + for key, value in expect.items(): + assert document[key] == value + for key, value in changes.items(): + if value is UNSET: + assert key not in document + assert model.to_json()["extra"] == 0 + + +@pytest.mark.parametrize( + ("changes", "at"), + [ + ({"dtype": "float32"}, ("dtype",)), + ({"chunks": [1, 1]}, ("chunks",)), + ({"dtype": "|b1"}, ("fill_value",)), + ], + ids=["dtype", "rank", "fill-no-longer-fits"], +) +def test_error_update_refuses_a_document_with_a_problem(changes: dict[str, Any], at: Loc) -> None: + """A change that makes a document with a problem is refused at the change, with the problem where it is: a dtype the family's fill value no longer fits is reported at the fill value.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ArrayMetadata(ARRAY).update(**changes) + assert raised.value.problems[0].loc == at + + +def test_with_context_and_refined_in_read_the_document_in_another_scope() -> None: + """`with_context` reads the document in any scope (a loss is allowed), `refined_in` only up the order: a scope that claims what this one left unclaimed is a gain, one that reads a name by another definition or by none is a `ScopeConflictError` naming where the name sits, and a gain that surfaces a problem is a `MetadataValidationError`.""" + unclaimed = ZarrV2ArrayMetadata({**ARRAY, "compressor": {"id": "zlib"}}, SMALL) + assert isinstance(unclaimed.compressor, UnclaimedField) + gained = unclaimed.refined_in(PRIVATE) + assert isinstance(gained.compressor, AcceptedField) + assert gained.compressor.definition is BARE_ZLIB + assert unclaimed.refines(unclaimed) + assert gained.refines(unclaimed) + assert not unclaimed.refines(gained) + with pytest.raises(ScopeConflictError) as conflict: + gained.refined_in(CORE_V2) + assert [c.loc for c in conflict.value.conflicts] == [("compressor",)] + lost = gained.with_context(SMALL) + assert isinstance(lost.compressor, UnclaimedField) + assert lost == unclaimed + assert gained.with_context(PRIVATE) == gained + with pytest.raises(MetadataValidationError): + ZarrV2ArrayMetadata({**ARRAY, "compressor": {"id": "zlib", "level": 1}}, SMALL).refined_in( + PRIVATE + ) + + +def test_refines_orders_models_by_information() -> None: + """`refines` holds when each field refines its counterpart and every other member is the same, the fill value as the more informed dtype spells it; a fill value that dtype refuses is no refinement; a value of another type refines nothing.""" + core = ZarrV2ArrayMetadata(ARRAY) + assert core.refines(ZarrV2ArrayMetadata({**ARRAY, "fill_value": 0.0})) + assert not core.refines(ZarrV2ArrayMetadata({**ARRAY, "shape": [8], "chunks": [2]})) + small = ZarrV2ArrayMetadata({**ARRAY, "dtype": " None: + """No model is invalid: the constructor raises with every problem the read finds.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ArrayMetadata({**ARRAY, "dtype": "float32", "order": "Q"}) + assert [p.loc for p in raised.value.problems] == [("dtype",), ("order",)] + + +def test_the_dataclass_machinery_is_gone() -> None: + """A v2 model is built only from a document: `dataclasses.replace` and `dataclasses.fields` do not apply, and the old `...Partial` names are gone.""" + import zarr_metadata + + assert not dataclasses.is_dataclass(ZarrV2ArrayMetadata) + assert not hasattr(zarr_metadata, "ZarrV2ArrayMetadataPartial") + assert "ZarrV2ArrayMetadataUpdate" in zarr_metadata.__all__ + + +GROUP: dict[str, Any] = {"zarr_format": 2, "attributes": {"g": [1]}} + + +@pytest.mark.parametrize( + ("document", "attributes"), + [ + (GROUP, {"g": (1,)}), + ({"zarr_format": 2}, UNSET), + ({"zarr_format": 2, "attributes": {}}, {}), + ], + ids=["attributes", "no-zattrs", "empty-zattrs"], +) +def test_a_group_is_its_document_read_in_its_scope( + document: dict[str, Any], attributes: object +) -> None: + """A v2 group model is its document and scope: `attributes` as the read refined them, `UNSET` when no `.zattrs` exists; it round-trips through its document, its store keys, pickle and copy, in any scope.""" + model = ZarrV2GroupMetadata(document, SMALL) + assert model.context is SMALL + assert model.attributes == attributes + assert model.to_json() == ZarrV2GroupMetadata.from_json(document).to_json() + assert ZarrV2GroupMetadata.from_key_value(model.to_key_value(), context=SMALL) == model + assert pickle.loads(pickle.dumps(model)) == model + assert copy.deepcopy(model) == model + assert dict(model.claims) == {} + + +def test_a_group_updates_moves_scope_and_refines_as_the_arrays_do() -> None: + """`update` reads the new attributes in the model's own scope and `UNSET` leaves them out; `with_context` and `refined_in` read the document in another scope and conflict with nothing, since a group holds no field; two groups are one when their attributes are written alike, and `refines` is that equality.""" + model = ZarrV2GroupMetadata(GROUP) + assert model.update(attributes={"h": 2}).attributes == {"h": 2} + assert model.update(attributes=UNSET).attributes is UNSET + assert model.refined_in(SMALL).context is SMALL + assert model.with_context(PRIVATE) == model + assert model == ZarrV2GroupMetadata({"zarr_format": 2, "attributes": {"g": (1,)}}) + assert model != ZarrV2GroupMetadata({"zarr_format": 2}) + assert model.refines(model) + assert not model.refines(ZarrV2GroupMetadata({"zarr_format": 2})) + assert not dataclasses.is_dataclass(ZarrV2GroupMetadata) + + +def test_error_a_group_document_with_a_problem_is_refused_at_construction() -> None: + """A group document with a problem -- a member the spec does not define, an attribute key that is not a string -- is refused with every problem.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV2GroupMetadata({"zarr_format": 2, "attributes": {1: "a"}, "extra": 1}) + assert sorted((p.loc, p.kind) for p in raised.value.problems) == [ + (("attributes",), "invalid_type"), + (("extra",), "unknown_key"), + ] + + +ZARRAY = {k: v for k, v in ARRAY.items() if k != "attributes"} +CONSOLIDATED: dict[str, Any] = { + "zarr_consolidated_format": 1, + "metadata": { + ".zgroup": {"zarr_format": 2}, + ".zattrs": {"root": True}, + "a/.zarray": ZARRAY, + "a/.zattrs": {"a": [1, 2]}, + "b/.zgroup": {"zarr_format": 2}, + "orphan/.zattrs": {"o": 1}, + }, +} + + +def _consolidated(**entries: object) -> dict[str, Any]: + return {**CONSOLIDATED, "metadata": {**CONSOLIDATED["metadata"], **entries}} + + +def test_consolidated_metadata_holds_its_entries_verbatim_and_each_node_as_a_model() -> None: + """`metadata` is the flat file-keyed map as written, refined, read-only; `nodes` is each `.zarray`/`.zgroup` entry merged with its sibling `.zattrs` as a model of the consolidated scope, keyed by node path (`""` for the root); a `.zattrs` with no sibling is kept and makes no node; the whole round-trips through its document, store keys, pickle and copy.""" + model = ZarrV2ConsolidatedMetadata(CONSOLIDATED, PRIVATE) + assert model.context is PRIVATE + assert model.metadata["a/.zattrs"] == {"a": (1, 2)} + assert set(model.nodes) == {"", "a", "b"} + assert model.nodes["a"] == ZarrV2ArrayMetadata(ARRAY, PRIVATE) + root = ZarrV2GroupMetadata({"zarr_format": 2, "attributes": {"root": True}}, PRIVATE) + assert model.nodes[""] == root + assert model.nodes["b"].attributes is UNSET + entries = cast("dict[str, Any]", model.to_json()["metadata"]) + assert entries["orphan/.zattrs"] == {"o": 1} + assert ZarrV2ConsolidatedMetadata.from_key_value(model.to_key_value(), context=PRIVATE) == model + assert pickle.loads(pickle.dumps(model)) == model + assert copy.deepcopy(model) == model + with pytest.raises(TypeError): + cast("dict[str, object]", model.metadata)["x"] = 1 + assert not dataclasses.is_dataclass(ZarrV2ConsolidatedMetadata) + + +def test_consolidated_metadata_is_equal_by_its_nodes_and_moves_scope_with_them() -> None: + """Two consolidated documents are one when each node means the same and the other entries are written alike; `with_context`/`refined_in` read every node in the new scope, a conflict located at the node's entry; `refines` holds when each node refines its counterpart.""" + spelled = _consolidated(**{"a/.zarray": {**ZARRAY, "fill_value": 0.0}}) + assert ZarrV2ConsolidatedMetadata(CONSOLIDATED) == ZarrV2ConsolidatedMetadata(spelled) + other = _consolidated(**{"orphan/.zattrs": {"o": 2}}) + assert ZarrV2ConsolidatedMetadata(CONSOLIDATED) != ZarrV2ConsolidatedMetadata(other) + small = ZarrV2ConsolidatedMetadata(CONSOLIDATED, SMALL) + assert isinstance(small.nodes["a"], ZarrV2ArrayMetadata) + assert isinstance(small.nodes["a"].compressor, UnclaimedField) + gained = small.refined_in(PRIVATE) + assert isinstance(gained.nodes["a"], ZarrV2ArrayMetadata) + assert isinstance(gained.nodes["a"].compressor, AcceptedField) + assert gained.refines(small) + assert not small.refines(gained) + with pytest.raises(ScopeConflictError) as conflict: + gained.refined_in(CORE_V2) + assert [c.loc for c in conflict.value.conflicts] == [("metadata", "a/.zarray", "compressor")] + assert gained.with_context(SMALL) == small + + +@pytest.mark.parametrize( + ("entries", "at"), + [ + ({"a/.zarray": {**ZARRAY, "dtype": "float32"}}, ("metadata", "a/.zarray", "dtype")), + ({"a/.zarray": ZARRAY, "a/.zattrs": {1: 2}}, ("metadata", "a/.zattrs")), + ({"a/.zarray": ZARRAY, "a/.zgroup": {"zarr_format": 2}}, ("metadata", "a/.zgroup")), + ({"a/.zgroup": {"zarr_format": 3}}, ("metadata", "a/.zgroup", "zarr_format")), + ({"a/.zarray": 3}, ("metadata", "a/.zarray")), + ], + ids=["array-dtype", "zattrs-key", "array-and-group", "group-format", "not-an-object"], +) +def test_error_a_node_entry_with_a_problem_is_refused_at_the_entry( + entries: dict[str, Any], at: Loc +) -> None: + """A `.zarray` or `.zgroup` entry is read as the document it is, in the consolidated scope; its problems sit under the entry, a `.zattrs`'s under its own entry; a path that is both an array and a group is a problem at the `.zgroup`.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ConsolidatedMetadata({"zarr_consolidated_format": 1, "metadata": entries}) + assert raised.value.problems[0].loc == at + + +def test_error_two_entries_for_one_node_file_are_refused() -> None: + """Two keys that name one file of one node -- `.zarray` and `/.zarray` -- are a problem at the second, since one would otherwise go unread; and a conflict is located at the key the document writes.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ConsolidatedMetadata( + {"zarr_consolidated_format": 1, "metadata": {".zarray": ZARRAY, "/.zarray": ZARRAY}} + ) + assert [(p.loc, p.kind) for p in raised.value.problems] == [ + (("metadata", "/.zarray"), "invalid_value") + ] + gained = ZarrV2ConsolidatedMetadata( + {"zarr_consolidated_format": 1, "metadata": {"/.zarray": ZARRAY}}, PRIVATE + ) + with pytest.raises(ScopeConflictError) as conflict: + gained.refined_in(CORE_V2) + assert [c.loc for c in conflict.value.conflicts] == [("metadata", "/.zarray", "compressor")] + + +def test_the_node_type_is_exported() -> None: + """`ZarrV2NodeMetadata`, the type of each node `nodes` holds, is exported beside the models, as `ZarrV3NodeMetadata` is.""" + from zarr_metadata import model + + assert "ZarrV2NodeMetadata" in model.__all__ + + +def test_error_a_second_key_for_a_node_file_hides_no_other_problem() -> None: + """A second key naming one file of one node is one problem among the document's: every node is still read, and each node's problems reported, so a user sees everything at once; a leading or repeated `/` does not make a second node.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ConsolidatedMetadata( + { + "zarr_consolidated_format": 1, + "metadata": { + "a/.zarray": ZARRAY, + "/a/.zarray": ZARRAY, + "b/.zgroup": {"zarr_format": 3}, + }, + } + ) + assert [p.loc for p in raised.value.problems] == [ + ("metadata", "/a/.zarray"), + ("metadata", "b/.zgroup", "zarr_format"), + ] + one = ZarrV2ConsolidatedMetadata( + {"zarr_consolidated_format": 1, "metadata": {"/a/.zarray": ZARRAY}} + ) + assert set( + ZarrV2ConsolidatedMetadata( + { + "zarr_consolidated_format": 1, + "metadata": {"x/.zgroup": {"zarr_format": 2}, "x//y/.zarray": ZARRAY}, + } + ).nodes + ) == {"x", "x/y"} + assert set(one.nodes) == {"a"} + assert one == ZarrV2ConsolidatedMetadata( + {"zarr_consolidated_format": 1, "metadata": {"a/.zarray": ZARRAY}} + ) + + +def test_error_a_key_with_a_dot_segment_is_no_node_s_path() -> None: + """A `.zmetadata` key whose path has a `.` or `..` segment names no node, as a leading or repeated `/` does not: it is a problem at the key, and not a second node beside the one it would name.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ConsolidatedMetadata( + { + "zarr_consolidated_format": 1, + "metadata": { + ".zgroup": {"zarr_format": 2}, + "./a/.zarray": ZARRAY, + "b/../a/.zarray": ZARRAY, + }, + } + ) + assert [(p.loc, p.kind) for p in raised.value.problems] == [ + (("metadata", "./a/.zarray"), "invalid_value"), + (("metadata", "b/../a/.zarray"), "invalid_value"), + ] + + +@pytest.mark.parametrize( + ("entries", "at", "kind"), + [ + ( + {".zgroup": {"zarr_format": 2}, "a/.zarray": ZARRAY, "a/b/.zarray": ZARRAY}, + "a/b/.zarray", + "invalid_value", + ), + ({".zarray": ZARRAY, "x/.zgroup": {"zarr_format": 2}}, "x/.zgroup", "invalid_value"), + ({".zgroup": {"zarr_format": 2}, "x/y/.zarray": ZARRAY}, "x/.zgroup", "missing_key"), + ], + ids=["below-an-array", "below-the-root-array", "missing-parent"], +) +def test_error_the_nodes_of_a_zmetadata_make_a_hierarchy( + entries: dict[str, Any], at: str, kind: str +) -> None: + """The nodes a `.zmetadata` holds make a hierarchy, as a v3 group's consolidated metadata does: a node below an array is a problem at its entry, and a group missing above a node a `missing_key` at the entry the group would have.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ConsolidatedMetadata({"zarr_consolidated_format": 1, "metadata": entries}) + assert [(p.loc, p.kind) for p in raised.value.problems] == [(("metadata", at), kind)] diff --git a/packages/zarr-metadata/tests/model/test_pydantic.py b/packages/zarr-metadata/tests/model/test_pydantic.py index 7951cac5c3..bccbe528e1 100644 --- a/packages/zarr-metadata/tests/model/test_pydantic.py +++ b/packages/zarr-metadata/tests/model/test_pydantic.py @@ -42,7 +42,9 @@ ) from zarr_metadata import JSONValue -from zarr_metadata.model import ZarrV3ArrayMetadata, ZarrV3NamedConfig +from zarr_metadata.model import ( + ZarrV3ArrayMetadata, +) # --- the integration (this is the example) ----------------------------------- @@ -169,8 +171,6 @@ def build_and_use() -> None: "ZarrV3ExtensionField": ZarrV3ExtensionField, "ZarrV3MetadataFieldJSON": ZarrV3MetadataFieldJSON, "ZarrV3ArrayMetadataJSON": ZarrV3ArrayMetadataJSON, - "ZarrV3NamedConfig": ZarrV3NamedConfig, - "ZarrV3MetadataField": ZarrV3NamedConfig, "UNSET": UNSET, }, ) diff --git a/packages/zarr-metadata/tests/model/test_pydantic_module.py b/packages/zarr-metadata/tests/model/test_pydantic_module.py index cb45310759..ca843a46a9 100644 --- a/packages/zarr-metadata/tests/model/test_pydantic_module.py +++ b/packages/zarr-metadata/tests/model/test_pydantic_module.py @@ -9,6 +9,7 @@ import math import warnings from collections.abc import Mapping +from typing import Any import pytest from jsonschema import Draft202012Validator @@ -17,13 +18,18 @@ import zarr_metadata.pydantic as zmp from zarr_metadata._common import JSONValue from zarr_metadata.model import ( + ValidationProblem, ZarrV2ArrayMetadata, ZarrV2ConsolidatedMetadata, ZarrV2GroupMetadata, ZarrV3ArrayMetadata, ZarrV3ConsolidatedMetadata, ZarrV3GroupMetadata, - ZarrV3NamedConfig, +) +from zarr_metadata.v3.codec.gzip import GZIP_CODEC +from zarr_metadata.v3.definition import ( + CORE, + AcceptedField, ) V3_ARRAY_DOC = dict(ZarrV3ArrayMetadata.create_default(shape=(4,)).to_json()) @@ -57,7 +63,6 @@ V2_CONSOLIDATED_DOC, id="consolidated-v2", ), - pytest.param(zmp.ZarrV3MetadataField, ZarrV3NamedConfig, {"name": "bytes"}, id="field-v3"), ] @@ -96,8 +101,16 @@ class Manifest(BaseModel): doc = dict(V3_ARRAY_DOC) del doc["chunk_key_encoding"] - with pytest.raises(ValidationError, match="chunk_key_encoding: missing required key"): + with pytest.raises(ValidationError) as raised: Manifest.model_validate({"metadata": doc}) + (error,) = raised.value.errors() + assert (error["type"], error["loc"], error["msg"]) == ( + "missing_key", + ("metadata", "chunk_key_encoding"), + "missing required key", + ) + # A missing key's input is the object missing it, as pydantic's own is. + assert error["input"] == doc def test_json_schema_generation() -> None: @@ -105,7 +118,6 @@ def test_json_schema_generation() -> None: class Manifest(BaseModel): metadata: zmp.ZarrV3ArrayMetadata - codec: zmp.ZarrV3MetadataField schema = Manifest.model_json_schema() metadata_schema = schema["$defs"]["ZarrV3ArrayMetadataJSON"] @@ -125,10 +137,6 @@ class Manifest(BaseModel): "title": "Zarr Format", "type": "integer", } - assert schema["properties"]["codec"]["anyOf"] == [ - {"type": "string"}, - {"$ref": "#/$defs/ZarrV3NamedConfigJSON"}, - ] def test_json_schema_generation_emits_no_warnings() -> None: @@ -140,7 +148,6 @@ def test_json_schema_generation_emits_no_warnings() -> None: zmp.ZarrV2GroupMetadata, zmp.ZarrV3ConsolidatedMetadata, zmp.ZarrV2ConsolidatedMetadata, - zmp.ZarrV3MetadataField, ) with warnings.catch_warnings(): @@ -153,9 +160,10 @@ def test_v2_recursive_structured_dtype_is_in_pydantic_schema() -> None: """The schema accepts nested structured dtypes supported by the v2 specification.""" doc = json.loads(json.dumps(V2_ARRAY_DOC)) doc["dtype"] = [["outer", [["inner", " N def test_metadata_field_schema_rejects_unknown_members() -> None: """Named-configuration envelopes are closed in both runtime and schema validation.""" - _assert_runtime_and_schema_reject( - zmp.ZarrV3MetadataField, - {"name": "example", "unexpected": 1}, - ) + doc = json.loads(json.dumps(V3_ARRAY_DOC)) + doc["codecs"] = [{"name": "bytes", "unexpected": 1}] + + _assert_runtime_and_schema_reject(zmp.ZarrV3ArrayMetadata, doc) @pytest.mark.parametrize( @@ -235,12 +243,12 @@ def test_v2_schema_rejects_unknown_document_members( def test_v2_array_schema_allows_unknown_document_members() -> None: - """The v2 array document is open ("SHOULD be ignored", https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L91-L92), in runtime and schema.""" + """The v2 array document is open ("SHOULD be ignored", https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v2/v2.0.rst#L91-L92), in runtime and schema; the member is kept as written.""" doc = json.loads(json.dumps(V2_ARRAY_DOC)) doc["unexpected"] = 1 adapter = TypeAdapter(zmp.ZarrV2ArrayMetadata) - assert "unexpected" not in adapter.validate_python(doc).to_json() + assert adapter.validate_python(doc).to_json()["unexpected"] == 1 assert Draft202012Validator(adapter.json_schema()).is_valid(doc) @@ -264,15 +272,6 @@ class Manifest(BaseModel): assert Manifest.model_validate_json(manifest.model_dump_json()) == manifest -def test_metadata_field_serializes_shorthand_and_false_object() -> None: - """The optional integration exposes the core model's canonical extension form.""" - adapter = TypeAdapter(zmp.ZarrV3MetadataField) - assert adapter.dump_python(adapter.validate_python({"name": "bytes"})) == "bytes" - assert adapter.dump_python( - adapter.validate_python({"name": "optional", "must_understand": False}) - ) == {"name": "optional", "must_understand": False} - - def test_core_package_does_not_import_pydantic() -> None: """Importing zarr_metadata (in a fresh interpreter) must not import pydantic: the integration is opt-in via zarr_metadata.pydantic.""" @@ -296,3 +295,190 @@ class Manifest(BaseModel): held = Manifest.model_validate_json(written).metadata.attributes["x"] assert isinstance(held, float) assert math.isnan(held) + + +# --- the scope a v3 field type reads in -------------------------------------- + +_WITH_ZSTD = ZarrV3ArrayMetadata.create_default( + shape=(4,), + codecs=({"name": "bytes"}, {"name": "zstd", "configuration": {"level": 3, "checksum": False}}), +).to_json() + + +@pytest.mark.parametrize( + ("context", "read"), + [ + (None, True), + (CORE, False), + ({zmp.CONTEXT_KEY: CORE}, False), + ({"another validator's": 1}, True), + ], + ids=["none", "a-scope", "a-mapping-holding-one", "a-mapping-holding-none"], +) +def test_a_v3_field_type_reads_in_the_scope_the_validation_context_holds( + context: object, read: bool +) -> None: + """As pydantic hands any validator its context; `zstd` is an extension, which `CORE` leaves unclaimed.""" + model = TypeAdapter(zmp.ZarrV3ArrayMetadata).validate_python(_WITH_ZSTD, context=context) + assert isinstance(model.codecs[1], AcceptedField) is read + + +def test_error_a_validation_context_holds_a_scope_that_is_not_one() -> None: + with pytest.raises(TypeError, match=zmp.CONTEXT_KEY): + TypeAdapter(zmp.ZarrV3ArrayMetadata).validate_python( + _WITH_ZSTD, context={zmp.CONTEXT_KEY: "CORE"} + ) + + +def test_each_problem_is_a_line_error_of_pydantic_s() -> None: + # As pydantic reports its own: one per problem, its type the problem's + # kind, at the problem's loc under the field's, with the input found + # there and what was expected; a message holding JSON is kept as it is. + class Manifest(BaseModel): + metadata: zmp.ZarrV3ArrayMetadata + + doc = { + **V3_ARRAY_DOC, + "fill_value": {"a": 1}, + "codecs": [ + {"name": "bytes", "configuration": {"endian": "little"}}, + {"name": "gzip", "configuration": {"level": 12}}, + ], + } + with pytest.raises(ValidationError) as raised: + Manifest.model_validate({"metadata": doc}) + errors = raised.value.errors() + assert [(e["type"], e["loc"], e["input"]) for e in errors] == [ + ("invalid_type", ("metadata", "fill_value"), {"a": 1}), + ("invalid_value", ("metadata", "codecs", 1, "configuration", "level"), 12), + ] + assert errors[0]["msg"] == 'expected an integer, got {"a": 1}' + assert errors[1].get("ctx") == {"ge": 0, "le": 9} + + +def test_a_message_holding_a_ctx_placeholder_is_reported_as_it_is() -> None: + # pydantic renders a message as a template of its ctx, with no escape: + # a document value written `{expected}` would be shown as the ctx's + # `expected`. Such a message rides whole in the ctx instead. + doc = {**V3_ARRAY_DOC, "zarr_format": "{expected}"} + with pytest.raises(ValidationError) as raised: + TypeAdapter(zmp.ZarrV3ArrayMetadata).validate_python(doc) + (error,) = raised.value.errors() + assert (error["type"], error["loc"], error["msg"]) == ( + "invalid_type", + ("zarr_format",), + 'expected 3, got "{expected}"', + ) + assert error.get("ctx") == {"expected": (3,), "message": 'expected 3, got "{expected}"'} + + +def test_a_ctx_member_named_message_yields_to_a_message_holding_a_placeholder() -> None: + # The carrier goes last, so what it carries is scanned for no key + # after it; a member of its name is kept when nothing rides. + ctx = {"message": "theirs", "expected": 3} + problem = ValidationProblem(("a",), 'got "{expected}"', "invalid_value", ctx=ctx) + error = ValidationError.from_exception_data("T", [zmp._line_error(problem, {"a": 1})]).errors()[ + 0 + ] + assert error["msg"] == 'got "{expected}"' + assert error.get("ctx") == {"expected": 3, "message": 'got "{expected}"'} + assert list(error.get("ctx", {})) == ["expected", "message"] + plain = ValidationProblem(("a",), "got 1", "invalid_value", ctx=ctx) + error = ValidationError.from_exception_data("T", [zmp._line_error(plain, {"a": 1})]).errors()[0] + assert (error["msg"], error.get("ctx")) == ("got 1", ctx) + + +def test_a_line_error_for_what_a_problem_could_not_hold_reports_what_sits_there() -> None: + # A problem holds no input for what is not JSON a reader walks -- a + # `AcceptedField` built by hand among the codecs -- and the line error reports + # that object, where a missing key's reports the object missing it. + smuggled = AcceptedField( + json="gzip", name="gzip", definition=GZIP_CODEC, configuration={"level": 1} + ) + doc = { + **V3_ARRAY_DOC, + "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}, smuggled], + } + with pytest.raises(ValidationError) as raised: + TypeAdapter(zmp.ZarrV3ArrayMetadata).validate_python(doc) + (error,) = raised.value.errors() + assert (error["type"], error["loc"]) == ("invalid_type", ("codecs", 1)) + assert error["input"] is smuggled + + +def test_the_pydantic_schema_refuses_a_null_consolidated_metadata_as_the_reader_does() -> None: + # The package publishes one verdict on `null` there: the reader's. + schema = TypeAdapter(zmp.ZarrV3GroupMetadata).json_schema() + document = {"zarr_format": 3, "node_type": "group", "consolidated_metadata": None} + assert not Draft202012Validator(schema).is_valid(document) + with pytest.raises(ValidationError): + TypeAdapter(zmp.ZarrV3GroupMetadata).validate_python(document) + + +def test_a_line_error_for_a_document_past_the_cap_renders() -> None: + # The depth problem holds no input; the line error reports the subtree + # at its loc, which pydantic renders itself, truncated in `str` and + # whole in `json`, as it renders any input. + deep: dict[str, object] = {} + for _ in range(2_000): + deep = {"a": deep} + with pytest.raises(ValidationError) as raised: + TypeAdapter(zmp.ZarrV3ArrayMetadata).validate_python({**V3_ARRAY_DOC, "attributes": deep}) + (error,) = raised.value.errors() + assert (error["type"], len(error["loc"])) == ("invalid_value", 256) + assert "nested deeper" in str(raised.value) + assert json.loads(raised.value.json())[0]["type"] == "invalid_value" + + +@pytest.mark.parametrize("field", ["Foo/bar", "", "r*", {"name": "Int8"}]) +def test_the_pydantic_schema_names_an_extension_as_the_reader_does(field: object) -> None: + # One verdict on a name, the reader's, in the schema pydantic generates too. + schema = TypeAdapter(zmp.ZarrV3ArrayMetadata).json_schema() + # As JSON has it, arrays as lists: `jsonschema` takes no tuple for one. + document: dict[str, Any] = json.loads(json.dumps(V3_ARRAY_DOC)) + assert Draft202012Validator(schema).is_valid(document) + named: dict[str, Any] = {**document, "data_type": field} + assert not Draft202012Validator(schema).is_valid(named) + with pytest.raises(ValidationError): + TypeAdapter(zmp.ZarrV3ArrayMetadata).validate_python(named) + + +def test_v2_field_types_read_in_the_validation_contexts_scope() -> None: + """The v2 field types read a document in the scope the validation context holds -- itself a `Context`, or its `zarr_metadata_context` item -- and in `CORE_V2` when it holds none, as the v3 field types read in theirs.""" + from zarr_metadata.v2.data_type.scalar import UINT_V2 + from zarr_metadata.v2.definition import CORE_V2, Context, UnclaimedField + + adapter = TypeAdapter(zmp.ZarrV2ArrayMetadata) + doc = json.loads(json.dumps(V2_ARRAY_DOC)) + assert adapter.validate_python(doc).context is CORE_V2 + small = Context.of(UINT_V2) + read = adapter.validate_python({**doc, "compressor": {"id": "zlib"}}, context=small) + assert read.context is small + assert isinstance(read.compressor, UnclaimedField) + held = adapter.validate_python(doc, context={"zarr_metadata_context_v2": small}) + assert held.context is small + group = TypeAdapter(zmp.ZarrV2GroupMetadata).validate_python(V2_GROUP_DOC, context=small) + assert group.context is small + + +def test_each_format_reads_in_its_own_context_key() -> None: + """A mapping context names each format's scope by its own key -- `zarr_metadata_context` for v3, `zarr_metadata_context_v2` for v2 -- so a v3 scope given for the v3 fields leaves the v2 fields in `CORE_V2`; a bare `Context` is the scope of every field type.""" + from zarr_metadata.v2.data_type.scalar import UINT_V2 + from zarr_metadata.v2.definition import CORE_V2, Context, UnclaimedField + from zarr_metadata.v3.definition import CORE_AND_EXTENSIONS + + adapter = TypeAdapter(zmp.ZarrV2ArrayMetadata) + doc = json.loads(json.dumps(V2_ARRAY_DOC)) + small = Context.of(UINT_V2) + v3_only = adapter.validate_python(doc, context={"zarr_metadata_context": CORE_AND_EXTENSIONS}) + assert v3_only.context is CORE_V2 + with pytest.raises(ValidationError): + adapter.validate_python( + {**doc, "fill_value": "garbage"}, context={"zarr_metadata_context": CORE_AND_EXTENSIONS} + ) + own = adapter.validate_python(doc, context={"zarr_metadata_context_v2": small}) + assert own.context is small + bare = adapter.validate_python({**doc, "compressor": {"id": "zlib"}}, context=small) + assert bare.context is small + assert isinstance(bare.compressor, UnclaimedField) + assert zmp.CONTEXT_KEY_V2 == "zarr_metadata_context_v2" diff --git a/packages/zarr-metadata/tests/model/test_read_array_metadata.py b/packages/zarr-metadata/tests/model/test_read_array_metadata.py new file mode 100644 index 0000000000..6c0555c585 --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_read_array_metadata.py @@ -0,0 +1,397 @@ +"""A v3 array document, read: each field as a scope read it, and its codecs as a pipeline. + +`read_array_metadata_v3` returns everything one read of a document finds: +each field with where it sits, the kind it was read as and what the scope +made of it, the chunks the codecs are handed, each codec with the chunk it +is handed, every problem, and the document's model when there is none. +""" + +from __future__ import annotations + +from typing import TYPE_CHECKING, Any, cast + +import pytest + +from zarr_metadata._json import arrays_to_tuples +from zarr_metadata.model import ( + UNSET, + ValidationProblem, + ZarrV3ArrayMetadata, + ZarrV3ArrayMetadataReading, + read_array_metadata_v3, + read_group_metadata_v3, +) +from zarr_metadata.v3._definition import ( + fields_of, + with_problems, +) +from zarr_metadata.v3.definition import ( + CORE, + CORE_AND_EXTENSIONS, + AcceptedField, + ChunkGridDefinition, + ChunkKeyEncodingDefinition, + CodecDefinition, + Context, + DataTypeDefinition, + Definition, + RefusedField, + StorageTransformerDefinition, + UnclaimedField, + resolve, +) + +if TYPE_CHECKING: + from collections.abc import Callable, Iterator + + from zarr_metadata.v3.definition import Lengths, Loc, ResolvedField + +LITTLE = {"name": "bytes", "configuration": {"endian": "little"}} +ZSTD = {"name": "zstd", "configuration": {"level": 1}} + +POINTS: list[tuple[Loc, type[Definition[Any]], type]] = [ + (("data_type",), DataTypeDefinition, AcceptedField), + (("chunk_grid",), ChunkGridDefinition, AcceptedField), + (("chunk_key_encoding",), ChunkKeyEncodingDefinition, AcceptedField), +] +"""A default document's single extension points, each read where it sits.""" + + +def _document(shape: tuple[int, ...] = (4,), **fields: object) -> dict[str, Any]: + """A default v3 array document of `shape` with `fields` in it, arrays as tuples as a reader's are.""" + document = {**ZarrV3ArrayMetadata.create_default(shape=shape).to_json(), **fields} + return cast("dict[str, Any]", arrays_to_tuples(document)) + + +def _codec(loc: Loc, variant: type = AcceptedField) -> tuple[Loc, type[Definition[Any]], type]: + return (loc, CodecDefinition, variant) + + +def _shard(chunk_shape: list[int], codecs: list[object] | None = None, **more: object) -> object: + configuration = { + "chunk_shape": chunk_shape, + "codecs": [LITTLE] if codecs is None else codecs, + "index_codecs": [LITTLE], + **more, + } + return {"name": "sharding_indexed", "configuration": configuration} + + +@pytest.mark.parametrize( + ("document", "context", "fields", "handed"), + [ + (_document(), CORE_AND_EXTENSIONS, [*POINTS, _codec(("codecs", 0))], [(frozenset({4}),)]), + # Each codec is handed the chunk the one before it hands on. + ( + _document( + (4, 6), codecs=[{"name": "transpose", "configuration": {"order": [1, 0]}}, "bytes"] + ), + CORE_AND_EXTENSIONS, + [*POINTS, _codec(("codecs", 0)), _codec(("codecs", 1))], + [(frozenset({4}), frozenset({6})), (frozenset({6}), frozenset({4}))], + ), + # The fields a shard holds come after it, where they sit in its + # configuration; read in the core spec's scope, an extension inside + # it is out of scope where it sits. + ( + _document( + (16, 16), + chunk_grid={"name": "regular", "configuration": {"chunk_shape": [8, 8]}}, + codecs=[_shard([4, 4], codecs=[LITTLE, ZSTD], index_codecs=[LITTLE, "crc32c"])], + ), + CORE, + [ + *POINTS, + _codec(("codecs", 0)), + _codec(("codecs", 0, "configuration", "codecs", 0)), + _codec(("codecs", 0, "configuration", "codecs", 1), UnclaimedField), + _codec(("codecs", 0, "configuration", "index_codecs", 0)), + _codec(("codecs", 0, "configuration", "index_codecs", 1)), + ], + [(frozenset({8}), frozenset({8}))], + ), + # A shard within a shard: each field, then the fields it holds. + ( + _document( + (16, 16), + chunk_grid={"name": "regular", "configuration": {"chunk_shape": [8, 8]}}, + codecs=[_shard([4, 4], codecs=[_shard([2, 2])])], + ), + CORE_AND_EXTENSIONS, + [ + *POINTS, + _codec(("codecs", 0)), + _codec(("codecs", 0, "configuration", "codecs", 0)), + _codec(("codecs", 0, "configuration", "codecs", 0, "configuration", "codecs", 0)), + _codec( + ("codecs", 0, "configuration", "codecs", 0, "configuration", "index_codecs", 0) + ), + _codec(("codecs", 0, "configuration", "index_codecs", 0)), + ], + [(frozenset({8}), frozenset({8}))], + ), + # A shard its check refuses keeps the fields it read inside it. + ( + _document( + (16, 16), + chunk_grid={"name": "regular", "configuration": {"chunk_shape": [8, 8]}}, + codecs=[_shard([4, 4], index_location="middle", codecs=[LITTLE, ZSTD])], + ), + CORE, + [ + *POINTS, + _codec(("codecs", 0), RefusedField), + _codec(("codecs", 0, "configuration", "codecs", 0)), + _codec(("codecs", 0, "configuration", "codecs", 1), UnclaimedField), + _codec(("codecs", 0, "configuration", "index_codecs", 0)), + ], + [(frozenset({8}), frozenset({8}))], + ), + # A struct's field types are data types, whether or not anything + # in scope claims them. + ( + _document( + data_type={ + "name": "struct", + "configuration": { + "fields": [ + {"name": "a", "data_type": "int8"}, + {"name": "b", "data_type": "acme.t"}, + ] + }, + }, + fill_value={"a": 0, "b": "anything"}, + codecs=[LITTLE], + ), + CORE_AND_EXTENSIONS, + [ + POINTS[0], + ( + ("data_type", "configuration", "fields", 0, "data_type"), + DataTypeDefinition, + AcceptedField, + ), + ( + ("data_type", "configuration", "fields", 1, "data_type"), + DataTypeDefinition, + UnclaimedField, + ), + *POINTS[1:], + _codec(("codecs", 0)), + ], + [(frozenset({4}),)], + ), + # A field its definition refuses, or nothing claims, is read as + # the kind it sits in; a grid nothing claims says no lengths, and + # a codec after the array -> bytes codec is handed bytes. + ( + _document( + chunk_grid={"name": "acme.grid", "configuration": {}}, + codecs=["bytes", {"name": "gzip", "configuration": {"level": 99}}], + storage_transformers=[{"name": "acme.transformer"}], + ), + CORE_AND_EXTENSIONS, + [ + POINTS[0], + (("chunk_grid",), ChunkGridDefinition, UnclaimedField), + POINTS[2], + _codec(("codecs", 0)), + _codec(("codecs", 1), RefusedField), + (("storage_transformers", 0), StorageTransformerDefinition, UnclaimedField), + ], + [(None,), None], + ), + # A data type that is not JSON is refused. + ( + _document(data_type=float("nan")), + CORE_AND_EXTENSIONS, + [ + (("data_type",), DataTypeDefinition, RefusedField), + *POINTS[1:], + _codec(("codecs", 0)), + ], + [(frozenset({4}),)], + ), + ], + ids=[ + "default", + "transposed", + "a-shard-holding-an-extension-in-the-core-scope", + "a-shard-within-a-shard", + "a-shard-its-check-refuses", + "a-struct-s-field-types", + "fields-read-as-nothing", + "a-data-type-that-is-not-json", + ], +) +def test_a_document_reads_as_each_field_where_it_sits_and_its_codecs_as_a_pipeline( + document: dict[str, Any], + context: Context, + fields: list[tuple[Loc, type[Definition[Any]], type]], + handed: list[Lengths | None], +) -> None: + reading = read_array_metadata_v3(document, context=context) + # A model only of a document with no problem. + assert (reading.metadata is None) is (reading.problems != ()) + assert [(loc, field.read_as, type(field)) for loc, field in reading.fields()] == fields + # The chunk's vocabulary for what nothing says is None, the reading's + # for a key the document lacks is `UNSET`. + if reading.data_type is UNSET: + assert reading.chunk.data_type is None + else: + assert reading.chunk.data_type is reading.data_type + assert reading.pipeline[0].incoming == reading.chunk + assert [None if s.incoming is None else s.incoming.lengths for s in reading.pipeline] == handed + + +_INNER_GZIP = {"name": "gzip", "configuration": {"level": 12}} +_WITH_PROBLEMS = _document( + (4, 4), + chunk_grid={"name": "regular", "configuration": {"chunk_shape": [4]}}, + fill_value="high", + codecs=[ + _shard([2, 2], [LITTLE, _INNER_GZIP]), + ZSTD, + {"name": "transpose", "configuration": {"order": [1, 0]}}, + ], +) +"""An array document with a problem in each place a field can have one, and one in no field.""" + + +def _array_fields( + document: object, +) -> Iterator[tuple[Loc, ResolvedField[Any], tuple[ValidationProblem, ...]]]: + reading = read_array_metadata_v3(document) + return with_problems(reading.fields(), reading.problems) + + +def _group_fields( + document: object, +) -> Iterator[tuple[Loc, ResolvedField[Any], tuple[ValidationProblem, ...]]]: + group = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": { + "kind": "inline", + "must_understand": False, + "metadata": {"a": document}, + }, + } + reading = read_group_metadata_v3(group) + return with_problems(reading.fields(), reading.problems) + + +def _field_fields( + document: object, +) -> Iterator[tuple[Loc, ResolvedField[Any], tuple[ValidationProblem, ...]]]: + resolved, problems = resolve( + cast("dict[str, Any]", document)["codecs"][0], + CodecDefinition, + CORE_AND_EXTENSIONS, + ("codecs", 0), + ) + return with_problems(fields_of(resolved, ("codecs", 0)), problems) + + +_INNER_LEVEL = ("codecs", 0, "configuration", "codecs", 1, "configuration", "level") +_EACH_FIELD_S = { + ("data_type",): [], + ("chunk_grid",): [("chunk_grid", "configuration", "chunk_shape")], + ("chunk_key_encoding",): [], + ("codecs", 0): [_INNER_LEVEL], + ("codecs", 0, "configuration", "codecs", 0): [], + ("codecs", 0, "configuration", "codecs", 1): [_INNER_LEVEL], + ("codecs", 0, "configuration", "index_codecs", 0): [], + ("codecs", 1): [], + ("codecs", 2): [("codecs", 2)], +} +"""Each field of `_WITH_PROBLEMS`, and where each of its problems is.""" +_IN_A_GROUP = ("consolidated_metadata", "metadata", "a") + + +@pytest.mark.parametrize( + ("fields", "expected"), + [ + (_array_fields, _EACH_FIELD_S), + ( + _group_fields, + { + (*_IN_A_GROUP, *loc): [(*_IN_A_GROUP, *problem) for problem in problems] + for loc, problems in _EACH_FIELD_S.items() + }, + ), + ( + _field_fields, + {loc: problems for loc, problems in _EACH_FIELD_S.items() if loc[:2] == ("codecs", 0)}, + ), + ], + ids=["an-array", "a-document-a-group-holds", "one-field"], +) +def test_each_field_comes_with_the_problems_located_in_it( + fields: Callable[ + [object], Iterator[tuple[Loc, ResolvedField[Any], tuple[ValidationProblem, ...]]] + ], + expected: dict[Loc, list[Loc]], +) -> None: + # Those it was read with and those the document found with it where + # it stands -- a transpose out of the pipeline's order, at its own + # place, and the chunk grid over the shape -- the fields it holds + # too: a shard with a bad inner codec has that problem as well. A + # problem in no field -- the fill value -- is in none's. + assert { + loc: [p.loc for p in problems] for loc, _, problems in fields(_WITH_PROBLEMS) + } == expected + + +def test_a_document_with_no_problem_reads_as_its_model_holding_the_fields_read() -> None: + # The fields the read made, not a second reading of them. + document = _document(codecs=[{"name": "transpose", "configuration": {"order": [0]}}, LITTLE]) + reading = read_array_metadata_v3(document) + model = reading.metadata + assert reading.problems == () + assert model is not None + assert model == ZarrV3ArrayMetadata.from_json(document) + assert model.data_type is reading.data_type + assert model.chunk_grid is reading.chunk_grid + assert model.chunk_key_encoding is reading.chunk_key_encoding + assert all( + codec is stage.codec for codec, stage in zip(model.codecs, reading.pipeline, strict=True) + ) + + +def test_error_a_value_that_is_not_a_mapping_reads_as_nothing() -> None: + reading = read_array_metadata_v3(["not", "a", "document"]) + not_an_object = ValidationProblem((), "expected an object", "invalid_type") + assert reading == ZarrV3ArrayMetadataReading(problems=(not_an_object,)) + assert list(reading.fields()) == [] + + +@pytest.mark.parametrize( + ("member", "attribute"), + [("codecs", "pipeline"), ("storage_transformers", "storage_transformers")], +) +def test_error_a_list_of_fields_that_is_not_a_list_reads_as_empty( + member: str, attribute: str +) -> None: + document = _document() + document[member] = "bytes" + reading = read_array_metadata_v3(document) + assert getattr(reading, attribute) == () + assert [(p.loc, p.kind) for p in reading.problems] == [((member,), "invalid_type")] + + +@pytest.mark.parametrize("member", ["data_type", "chunk_grid", "chunk_key_encoding"]) +def test_error_a_document_without_a_field_reads_it_as_unset(member: str) -> None: + document = _document() + del document[member] + reading = read_array_metadata_v3(document) + assert getattr(reading, member) is UNSET + assert (member,) not in [loc for loc, _ in reading.fields()] + assert ValidationProblem((member,), "missing required key", "missing_key") in reading.problems + + +def test_error_a_document_without_a_shape_hands_its_codecs_chunks_of_no_known_rank() -> None: + document = _document() + del document["shape"] + reading = read_array_metadata_v3(document) + assert reading.chunk.lengths is None diff --git a/packages/zarr-metadata/tests/model/test_read_array_metadata_v2.py b/packages/zarr-metadata/tests/model/test_read_array_metadata_v2.py new file mode 100644 index 0000000000..dcd671a3b5 --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_read_array_metadata_v2.py @@ -0,0 +1,125 @@ +"""A v2 array document read once: each field as the scope read it, every problem, and the model.""" + +from __future__ import annotations + +from typing import Any + +import pytest + +from zarr_metadata._sentinel import UNSET +from zarr_metadata.model import ( + ZarrV2ArrayMetadata, + is_array_metadata_v2, + parse_array_metadata_v2, + read_array_metadata_v2, + validate_array_metadata_v2, +) +from zarr_metadata.v2.codec.compression import ZLIB_V2 +from zarr_metadata.v2.data_type.scalar import FLOAT_V2 +from zarr_metadata.v2.definition import ( + CORE_V2, + AcceptedField, + Context, + RefusedField, + UnclaimedField, + ZarrV2CodecDefinition, +) +from zarr_metadata.v3.definition import ( + EmptyConfiguration, +) + +Loc = tuple[str | int, ...] +BASE: dict[str, Any] = dict(ZarrV2ArrayMetadata.create_default(shape=(4,)).to_json()) +PRIVATE = Context.of(FLOAT_V2, ZLIB_V2) + + +@pytest.mark.parametrize( + ("changes", "context", "dtype", "compressor", "filters", "locs"), + [ + ({}, None, AcceptedField, None, None, [("dtype",)]), + ( + { + "dtype": " None: + """`read_array_metadata_v2` reads the document once in the scope given (`CORE_V2` by default): `dtype`, `compressor` and `filters` are each `AcceptedField`, `UnclaimedField` or None as written, `fields()` walks them in document order with a struct's record types after it, and a document with no problem has none.""" + reading = read_array_metadata_v2({**BASE, **changes}, context=context) + assert reading.problems == () + assert type(reading.dtype) is dtype + assert (None if reading.compressor is None else type(reading.compressor)) == compressor + assert reading.filters is not UNSET + found = None if reading.filters is None else tuple(type(f) for f in reading.filters) + assert found == filters + assert [loc for loc, _ in reading.fields()] == locs + + +@pytest.mark.parametrize( + ("value", "problems"), + [ + ({**BASE, "dtype": "float32"}, [(("dtype",), "invalid_value")]), + ( + {**BASE, "compressor": {"id": "zlib", "level": 10}}, + [(("compressor", "level"), "invalid_value")], + ), + ({**BASE, "fill_value": "NaN"}, [(("fill_value",), "invalid_type")]), + (3, [((), "invalid_type")]), + ], + ids=["dtype", "compressor", "fill", "not-an-object"], +) +def test_error_a_reading_reports_what_the_validator_reports( + value: object, problems: list[tuple[Loc, str]] +) -> None: + """A reading of a document with a problem holds every problem `validate_array_metadata_v2` finds, a `RefusedField` field where one was refused, and no model.""" + reading = read_array_metadata_v2(value) + assert [(p.loc, p.kind) for p in reading.problems] == problems + assert reading.metadata is None + assert [(p.loc, p.kind) for p in validate_array_metadata_v2(value)] == problems + if isinstance(value, dict) and "dtype" in problems[0][0]: + assert isinstance(reading.dtype, RefusedField) + + +def test_the_v2_readers_read_in_the_scope_given() -> None: + """`validate_array_metadata_v2`, `is_array_metadata_v2` and `parse_array_metadata_v2` take a scope: in one that does not claim `|u1`, the default document still validates (an unclaimed dtype is left unjudged), and in one where a private `zlib` takes no level, a written level is a problem.""" + assert validate_array_metadata_v2(BASE, context=PRIVATE) == () + assert is_array_metadata_v2(BASE, context=PRIVATE) + assert parse_array_metadata_v2(BASE, context=PRIVATE)["dtype"] == "|u1" + bare = ZarrV2CodecDefinition(name="zlib", configuration=EmptyConfiguration) + doc = {**BASE, "compressor": {"id": "zlib", "level": 1}} + assert validate_array_metadata_v2(doc) == () + found = validate_array_metadata_v2(doc, context=CORE_V2.extended_with(bare)) + assert [p.loc for p in found] == [("compressor", "level")] diff --git a/packages/zarr-metadata/tests/model/test_refine_json.py b/packages/zarr-metadata/tests/model/test_refine_json.py index 5ce759d40a..d5594d4ebd 100644 --- a/packages/zarr-metadata/tests/model/test_refine_json.py +++ b/packages/zarr-metadata/tests/model/test_refine_json.py @@ -4,11 +4,11 @@ import math from collections import OrderedDict -from typing import TYPE_CHECKING +from typing import TYPE_CHECKING, cast import pytest -from zarr_metadata._json import refine_json, refine_user_data +from zarr_metadata._json import JSON_DEPTH, refine_json, refine_user_data, shown if TYPE_CHECKING: from collections.abc import Callable @@ -93,11 +93,126 @@ def test_error_every_leaf_that_is_not_json_is_reported() -> None: def test_a_value_nested_hundreds_deep_is_read() -> None: - # One frame per level of nesting: as deep as the interpreter goes, less - # what the test runner's own frames take. + # One frame per level of nesting, up to `JSON_DEPTH` of them. deep: dict[str, object] = {} - for _ in range(600): + for _ in range(JSON_DEPTH - 1): deep = {"k": deep} refined, problems = refine_json(deep) assert problems == () assert refined is not None + + +def test_error_a_value_nested_deeper_than_a_reader_walks() -> None: + # The level past the last is the problem, wherever it sits, so no + # document takes a reader past what the interpreter allows. + deep: dict[str, object] = {} + for _ in range(JSON_DEPTH + 40): + deep = {"k": deep} + refined, problems = refine_json({"a": [deep]}) + assert refined is None + assert [(len(p.loc), p.kind, p.message) for p in problems] == [ + (JSON_DEPTH, "invalid_value", f"nested deeper than the {JSON_DEPTH} levels a reader walks") + ] + assert problems[0].loc[:2] == ("a", 0) + + +def _nested(depth: int, innermost: object) -> list[object]: + """`innermost` inside `depth` arrays, each holding the next: `innermost` sits `depth` levels down.""" + value: object = innermost + for _ in range(depth): + value = [value] + return cast("list[object]", value) + + +@pytest.mark.parametrize( + ("value", "text"), + [ + (None, "null"), + (True, "true"), + ("C", '"C"'), + ((1, (2,)), "[1, [2]]"), + ({"a": None}, '{"a": null}'), + (float("nan"), "NaN"), + ({1: 2}, "{1: 2}"), + (_nested(JSON_DEPTH, []), "a value nested too deep to show"), + (_nested(JSON_DEPTH, {1}), "[" * JSON_DEPTH + "{1}" + "]" * JSON_DEPTH), + ], + ids=[ + "null", + "true", + "string", + "array", + "object", + "non-finite", + "not-json", + "past-the-levels", + "not-json-at-the-last-level", + ], +) +def test_a_value_is_shown_as_the_json_a_document_writes(value: object, text: str) -> None: + """A problem's message shows a value as its JSON, by its repr when it is not JSON, and says so when it nests past the levels a reader walks: a container there, not a non-JSON value sitting on the last level, which is shown.""" + assert shown(value) == text + + +def test_bytes_past_the_levels_a_reader_walks_are_not_json_rather_than_nested() -> None: + # `bytes` is a sequence to Python and no container to JSON, wherever + # it sits: past the cap it is still what it is, not nesting. + deep: list[object] = [b""] + for _ in range(JSON_DEPTH - 1): + deep = [deep] + _, problems = refine_json(deep) + assert [(len(p.loc), p.kind, p.message[:31]) for p in problems] == [ + (JSON_DEPTH, "invalid_type", "not a JSON-serializable value: ") + ] + + +def test_error_an_integer_of_more_digits_than_json_text_holds_is_a_problem( + interpreter_writes_4300_digits: None, +) -> None: + """An integer the interpreter will not convert to text, as `sys.get_int_max_str_digits` bounds one, is an `invalid_value` problem where it sits, so no reader raises on it: the document could not be written back, and `validate_*` and `read_*` agree.""" + from zarr_metadata.model import ( + ZarrV3ArrayMetadata, + read_array_metadata_v3, + validate_array_metadata_v3, + ) + + refined, problems = refine_json({"x": [10**5000]}) + assert refined is None + assert [(p.loc, p.kind) for p in problems] == [(("x", 0), "invalid_value")] + assert "digits" in problems[0].message + document = {**ZarrV3ArrayMetadata.create_default(shape=(4,)).to_json()} + document["attributes"] = {"x": 10**5000} + found = validate_array_metadata_v3(document) + assert [(p.loc, p.kind) for p in found] == [(("attributes", "x"), "invalid_value")] + assert read_array_metadata_v3(document).problems == found + shaped = {**document, "attributes": {}, "shape": [10**5000]} + assert [p.loc for p in validate_array_metadata_v3(shaped)] == [("shape", 0)] + + +def test_a_str_or_int_subclass_is_refined_to_the_json_type_it_is() -> None: + """A `StrEnum` member, an `IntEnum` member, or any `str`, `int` or `float` subclass, is refined to the plain value JSON writes for it, so a `Literal` and a tag read it as the value, while `bool` stays apart from `int`.""" + import enum + + from zarr_metadata.model import ZarrV3GroupMetadata + + class NodeType(enum.StrEnum): + GROUP = "group" + + class Format(enum.IntEnum): + THREE = 3 + + class Score(float): + pass + + refined, problems = refine_json( + {"a": NodeType.GROUP, "b": Format.THREE, "c": Score(1.5), "d": True} + ) + assert problems == () + assert refined == {"a": "group", "b": 3, "c": 1.5, "d": True} + assert isinstance(refined, dict) + assert all(type(refined[key]) in (str, int, float, bool) for key in refined) + assert type(refined["d"]) is bool + model = ZarrV3GroupMetadata.from_json( + {"zarr_format": Format.THREE, "node_type": NodeType.GROUP} + ) + assert model.to_json() == {"zarr_format": 3, "node_type": "group"} diff --git a/packages/zarr-metadata/tests/model/test_repair.py b/packages/zarr-metadata/tests/model/test_repair.py new file mode 100644 index 0000000000..76432baa90 --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_repair.py @@ -0,0 +1,229 @@ +"""Repairing v3 metadata a known writer bug made invalid, before a strict read.""" + +from __future__ import annotations + +import copy +from typing import Any + +import pytest + +from zarr_metadata.model import ( + MetadataValidationError, + ZarrV2ConsolidatedMetadata, + ZarrV3ArrayMetadata, + ZarrV3GroupMetadata, + read_node_metadata_v3, + read_repaired_consolidated_metadata_v2, + read_repaired_node_metadata_v3, + repair_consolidated_metadata_v2, + repair_node_metadata_v3, +) + +ARRAY: dict[str, Any] = dict(ZarrV3ArrayMetadata.create_default(shape=(0, 3)).to_json()) +GROUP: dict[str, Any] = {"zarr_format": 3, "node_type": "group"} + + +def _grid(*chunk_shape: int) -> dict[str, Any]: + return {"name": "regular", "configuration": {"chunk_shape": chunk_shape}} + + +def _consolidated(**metadata: object) -> dict[str, Any]: + envelope = {"kind": "inline", "must_understand": False, "metadata": metadata} + return {**GROUP, "consolidated_metadata": envelope} + + +ZERO_CHUNK = {**ARRAY, "chunk_grid": _grid(0, 3)} +NULL_CONSOLIDATED = {**GROUP, "consolidated_metadata": None} + + +@pytest.mark.parametrize( + ("value", "repaired", "repairs"), + [ + ( + ZERO_CHUNK, + {**ARRAY, "chunk_grid": _grid(1, 3)}, + [(("chunk_grid", "configuration", "chunk_shape", 0), "zero_chunk_length")], + ), + ( + {**ARRAY, "shape": (0, 0), "chunk_grid": _grid(0, 0)}, + {**ARRAY, "shape": (0, 0), "chunk_grid": _grid(1, 1)}, + [ + (("chunk_grid", "configuration", "chunk_shape", 0), "zero_chunk_length"), + (("chunk_grid", "configuration", "chunk_shape", 1), "zero_chunk_length"), + ], + ), + (NULL_CONSOLIDATED, GROUP, [(("consolidated_metadata",), "null_consolidated_metadata")]), + ( + _consolidated(a=ZERO_CHUNK, b=NULL_CONSOLIDATED), + _consolidated(a={**ARRAY, "chunk_grid": _grid(1, 3)}, b=GROUP), + [ + ( + ("consolidated_metadata", "metadata", "a", "chunk_grid", "configuration") + + ("chunk_shape", 0), + "zero_chunk_length", + ), + ( + ("consolidated_metadata", "metadata", "b", "consolidated_metadata"), + "null_consolidated_metadata", + ), + ], + ), + # No writer's bug: left for the strict read to judge. + (ARRAY, ARRAY, []), + ({**ARRAY, "chunk_grid": _grid(3, 0)}, {**ARRAY, "chunk_grid": _grid(3, 0)}, []), + ({**ARRAY, "chunk_grid": _grid(0)}, {**ARRAY, "chunk_grid": _grid(0)}, []), + ( + {key: item for key, item in ZERO_CHUNK.items() if key != "data_type"}, + {key: item for key, item in ARRAY.items() if key != "data_type"} + | {"chunk_grid": _grid(1, 3)}, + [(("chunk_grid", "configuration", "chunk_shape", 0), "zero_chunk_length")], + ), + ({**GROUP, "consolidated_metadata": 0}, {**GROUP, "consolidated_metadata": 0}, []), + ([ZERO_CHUNK], [ZERO_CHUNK], []), + ], + ids=[ + "zero-chunk", + "zero-chunks", + "null-consolidated", + "inside-consolidated", + "valid", + "zero-chunk-on-a-full-dimension", + "zero-chunk-of-another-rank", + "zero-chunk-without-a-data-type", + "consolidated-of-another-type", + "not-a-document", + ], +) +def test_repair_undoes_each_known_writer_bug( + value: object, repaired: object, repairs: list[tuple[tuple[str | int, ...], str]] +) -> None: + """Each known writer bug in a document, its consolidated documents' too, is undone and said where; anything else is left as it is, for the strict read to report, and the document handed in is not changed.""" + before = copy.deepcopy(value) + document, made = repair_node_metadata_v3(value) + assert document == repaired + assert [(repair.loc, repair.kind) for repair in made] == repairs + assert value == before + if len(repairs) == 0: + assert document is value + + +@pytest.mark.parametrize( + ("value", "model"), + [ + (ZERO_CHUNK, ZarrV3ArrayMetadata.from_json({**ARRAY, "chunk_grid": _grid(1, 3)})), + (NULL_CONSOLIDATED, ZarrV3GroupMetadata.from_json(GROUP)), + (ARRAY, ZarrV3ArrayMetadata.from_json(ARRAY)), + ], + ids=["zero-chunk", "null-consolidated", "valid"], +) +def test_read_repaired_reads_what_the_strict_read_refuses(value: object, model: object) -> None: + """A document a writer bug made invalid is refused by `read_node_metadata_v3` and read by `read_repaired_node_metadata_v3`, as the strict read reads the repaired document; a valid one reads the same either way.""" + repaired = read_repaired_node_metadata_v3(value) + assert repaired.reading.problems == () + assert repaired.reading.metadata == model + assert (len(read_node_metadata_v3(value).problems) == 0) is (len(repaired.repairs) == 0) + + +def test_read_repaired_reports_what_no_repair_applies_to() -> None: + """A problem no repair undoes is reported as the strict read reports it, beside the repairs that were made.""" + value = {key: item for key, item in ZERO_CHUNK.items() if key != "data_type"} + repaired = read_repaired_node_metadata_v3(value) + assert [repair.kind for repair in repaired.repairs] == ["zero_chunk_length"] + assert [(problem.loc, problem.kind) for problem in repaired.reading.problems] == [ + (("data_type",), "missing_key") + ] + assert repaired.reading.metadata is None + + +# --- v2 --------------------------------------------------------------------- + +ZGROUP_FROM_ZARR3: dict[str, Any] = { + "zarr_format": 2, + "consolidated_metadata": {"metadata": {}, "must_understand": False, "kind": "inline"}, +} +ZMETADATA_FROM_ZARR3: dict[str, Any] = { + "zarr_consolidated_format": 1, + "metadata": { + ".zgroup": {"zarr_format": 2}, + "b/.zgroup": ZGROUP_FROM_ZARR3, + "b/.zattrs": {"x": 1}, + }, +} +ZMETADATA_CLEAN: dict[str, Any] = { + "zarr_consolidated_format": 1, + "metadata": { + ".zgroup": {"zarr_format": 2}, + "b/.zgroup": {"zarr_format": 2}, + "b/.zattrs": {"x": 1}, + }, +} + + +@pytest.mark.parametrize( + ("value", "repaired", "repairs"), + [ + ( + ZMETADATA_FROM_ZARR3, + ZMETADATA_CLEAN, + [ + ( + ("metadata", "b/.zgroup", "consolidated_metadata"), + "consolidated_metadata_in_zgroup_entry", + ) + ], + ), + (ZMETADATA_CLEAN, ZMETADATA_CLEAN, []), + # The root's .zgroup: no zarr-python version writes the member there. + ( + {"zarr_consolidated_format": 1, "metadata": {".zgroup": ZGROUP_FROM_ZARR3}}, + {"zarr_consolidated_format": 1, "metadata": {".zgroup": ZGROUP_FROM_ZARR3}}, + [], + ), + ( + { + "zarr_consolidated_format": 1, + "metadata": {"b/.zgroup": {"zarr_format": 2, "consolidated_metadata": 3}}, + }, + { + "zarr_consolidated_format": 1, + "metadata": {"b/.zgroup": {"zarr_format": 2, "consolidated_metadata": 3}}, + }, + [], + ), + (3, 3, []), + ], + ids=["zarr-3-zgroup-entry", "clean", "root", "not-the-bug", "not-a-document"], +) +def test_repair_v2_undoes_the_consolidated_metadata_zarr_3_writes_into_a_zgroup_entry( + value: object, repaired: object, repairs: list[tuple[tuple[str | int, ...], str]] +) -> None: + """zarr-python 3.x writes a `consolidated_metadata` member into each non-root `.zgroup` entry of a `.zmetadata`, which the v2 group document does not take; `repair_consolidated_metadata_v2` removes it and says where; anything else is left as it is, and the document handed in is not changed.""" + before = copy.deepcopy(value) + document, made = repair_consolidated_metadata_v2(value) + assert document == repaired + assert [(repair.loc, repair.kind) for repair in made] == repairs + assert value == before + if len(repairs) == 0: + assert document is value + + +def test_read_repaired_consolidated_v2_reads_what_the_strict_read_refuses() -> None: + """The strict `ZarrV2ConsolidatedMetadata` refuses the zarr-python 3.x `.zmetadata` at the entry's member; `read_repaired_consolidated_metadata_v2` reads the repaired document, holds its model, and lists the repairs.""" + with pytest.raises(MetadataValidationError) as raised: + ZarrV2ConsolidatedMetadata(ZMETADATA_FROM_ZARR3) + assert [p.loc for p in raised.value.problems] == [ + ("metadata", "b/.zgroup", "consolidated_metadata") + ] + repaired = read_repaired_consolidated_metadata_v2(ZMETADATA_FROM_ZARR3) + assert repaired.problems == () + assert repaired.metadata == ZarrV2ConsolidatedMetadata(ZMETADATA_CLEAN) + assert [r.kind for r in repaired.repairs] == ["consolidated_metadata_in_zgroup_entry"] + broken = read_repaired_consolidated_metadata_v2( + { + **ZMETADATA_FROM_ZARR3, + "metadata": {**ZMETADATA_FROM_ZARR3["metadata"], "b/.zattrs": {1: "x"}}, + } + ) + assert broken.metadata is None + assert [p.loc for p in broken.problems] == [("metadata", "b/.zattrs")] + assert [r.kind for r in broken.repairs] == ["consolidated_metadata_in_zgroup_entry"] diff --git a/packages/zarr-metadata/tests/model/test_scope_threading.py b/packages/zarr-metadata/tests/model/test_scope_threading.py new file mode 100644 index 0000000000..5955e745ed --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_scope_threading.py @@ -0,0 +1,139 @@ +"""Every entry point reads in the scope it is given, and in `CORE_AND_EXTENSIONS` when given none.""" + +from __future__ import annotations + +import json +from typing import TYPE_CHECKING, Any + +import pytest + +import zarr_metadata.model as zm +from zarr_metadata.model import ( + ZarrV3ArrayMetadata, +) +from zarr_metadata.v3.definition import ( + CORE_AND_EXTENSIONS, + Context, +) + +if TYPE_CHECKING: + from collections.abc import Callable + +EMPTY = Context.of() +BASE: dict[str, Any] = dict(ZarrV3ArrayMetadata.create_default(shape=(4,)).to_json()) +BAD: dict[str, Any] = { + **BASE, + "codecs": (*BASE["codecs"], {"name": "gzip", "configuration": {"level": 12}}), +} +"""An array the default scope refuses -- a gzip level of 12 -- and an empty scope leaves unjudged.""" +GROUP: dict[str, Any] = { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": {"kind": "inline", "must_understand": False, "metadata": {"a": BAD}}, +} +STORE_A = {"zarr.json": json.dumps(BAD).encode()} +STORE_G = {"zarr.json": json.dumps(GROUP).encode()} + + +def _accepts(read: Callable[..., object]) -> Callable[[Context | None], bool]: + """Whether `read`, given `context`, finds nothing wrong: True for a model, a reading with a model, an empty problem tuple, or a guard saying yes.""" + + def accepted(context: Context | None) -> bool: + try: + found = read(context=context) + except zm.MetadataValidationError: + return False + if isinstance(found, tuple): + return len(found) == 0 + if isinstance(found, bool): + return found + reading = getattr(found, "reading", found) + metadata = getattr(reading, "metadata", found) + return metadata is not None + + return accepted + + +ENTRY_POINTS: dict[str, Callable[..., object]] = { + "validate_array_metadata_v3": lambda context: zm.validate_array_metadata_v3( + BAD, context=context + ), + "is_array_metadata_v3": lambda context: zm.is_array_metadata_v3( + zm.parse_array_metadata_v3(BAD, context=EMPTY), context=context + ), + "parse_array_metadata_v3": lambda context: zm.parse_array_metadata_v3(BAD, context=context), + "read_array_metadata_v3": lambda context: zm.read_array_metadata_v3(BAD, context=context), + "ZarrV3ArrayMetadata": lambda context: zm.ZarrV3ArrayMetadata(BAD, context=context), + "ZarrV3ArrayMetadata.from_json": lambda context: zm.ZarrV3ArrayMetadata.from_json( + BAD, context=context + ), + "ZarrV3ArrayMetadata.from_key_value": lambda context: zm.ZarrV3ArrayMetadata.from_key_value( + STORE_A, context=context + ), + "ZarrV3ArrayMetadata.create_default": lambda context: zm.ZarrV3ArrayMetadata.create_default( + shape=(4,), codecs=BAD["codecs"], context=context + ), + "validate_group_metadata_v3": lambda context: zm.validate_group_metadata_v3( + GROUP, context=context + ), + "is_group_metadata_v3": lambda context: zm.is_group_metadata_v3( + zm.parse_group_metadata_v3(GROUP, context=EMPTY), context=context + ), + "parse_group_metadata_v3": lambda context: zm.parse_group_metadata_v3(GROUP, context=context), + "read_group_metadata_v3": lambda context: zm.read_group_metadata_v3(GROUP, context=context), + "ZarrV3GroupMetadata": lambda context: zm.ZarrV3GroupMetadata(GROUP, context=context), + "ZarrV3GroupMetadata.from_json": lambda context: zm.ZarrV3GroupMetadata.from_json( + GROUP, context=context + ), + "ZarrV3GroupMetadata.from_key_value": lambda context: zm.ZarrV3GroupMetadata.from_key_value( + STORE_G, context=context + ), + "ZarrV3GroupMetadata.create_default": lambda context: zm.ZarrV3GroupMetadata.create_default( + consolidated_metadata=GROUP["consolidated_metadata"], context=context + ), + "ZarrV3ConsolidatedMetadata": lambda context: zm.ZarrV3ConsolidatedMetadata( + GROUP["consolidated_metadata"], context=context + ), + "ZarrV3ConsolidatedMetadata.from_json": lambda context: zm.ZarrV3ConsolidatedMetadata.from_json( + GROUP["consolidated_metadata"], context=context + ), + "read_node_metadata_v3": lambda context: zm.read_node_metadata_v3(GROUP, context=context), + "validate_node_metadata_v3": lambda context: zm.validate_node_metadata_v3( + GROUP, context=context + ), + "node_metadata_from_json_v3": lambda context: zm.node_metadata_from_json_v3( + GROUP, context=context + ), + "node_metadata_from_key_value_v3": lambda context: zm.node_metadata_from_key_value_v3( + STORE_G, context=context + ), + "read_repaired_node_metadata_v3": lambda context: zm.read_repaired_node_metadata_v3( + GROUP, context=context + ), +} + + +@pytest.mark.parametrize("read", ENTRY_POINTS.values(), ids=ENTRY_POINTS.keys()) +def test_every_entry_point_reads_in_the_scope_it_is_given(read: Callable[..., object]) -> None: + """Each entry point refuses a gzip level of 12 in the default scope, which `None` names too, and accepts it in a scope that leaves gzip unclaimed: the scope it is given is the scope it reads in.""" + accepted = _accepts(read) + assert accepted(EMPTY) is True + assert accepted(CORE_AND_EXTENSIONS) is False + assert accepted(None) is False + + +def test_the_json_schema_is_written_in_the_scope_it_is_given() -> None: + """`node_metadata_json_schema_v3` takes `None` for the default scope, and writes a different schema for an empty one.""" + assert zm.node_metadata_json_schema_v3(context=None) == zm.node_metadata_json_schema_v3() + assert zm.node_metadata_json_schema_v3(context=EMPTY) != zm.node_metadata_json_schema_v3() + + +@pytest.mark.parametrize("read", ENTRY_POINTS.values(), ids=ENTRY_POINTS.keys()) +def test_error_every_entry_point_refuses_a_scope_of_another_format( + read: Callable[..., object], +) -> None: + """A v3 entry point given a v2 scope raises `TypeError`: a scope reads documents of one format, and a v2 scope claims nothing a v3 document writes.""" + from zarr_metadata.v2.definition import CORE_V2 + + with pytest.raises(TypeError, match="format"): + read(context=CORE_V2) diff --git a/packages/zarr-metadata/tests/model/test_scope_threading_v2.py b/packages/zarr-metadata/tests/model/test_scope_threading_v2.py new file mode 100644 index 0000000000..f4f470458e --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_scope_threading_v2.py @@ -0,0 +1,131 @@ +"""Every v2 entry point reads in the scope it is given, `CORE_V2` when given none, and refuses a scope of another format.""" + +from __future__ import annotations + +import json +from typing import TYPE_CHECKING, Any + +import pytest + +import zarr_metadata.model as zm +from zarr_metadata.model import ZarrV2ArrayMetadata +from zarr_metadata.v2.definition import CORE_V2, Context, resolve_codec_v2, resolve_dtype_v2 +from zarr_metadata.v3.definition import CORE_AND_EXTENSIONS + +if TYPE_CHECKING: + from collections.abc import Callable + +EMPTY = Context.of() +BASE: dict[str, Any] = dict(ZarrV2ArrayMetadata.create_default(shape=(4,), chunks=(2,)).to_json()) +BAD: dict[str, Any] = {**BASE, "compressor": {"id": "gzip", "level": "x"}} +"""An array the v2 scope refuses -- a gzip level that is no integer -- and an empty scope leaves unjudged.""" +GROUP: dict[str, Any] = {"zarr_format": 2} +CONSOLIDATED: dict[str, Any] = { + "zarr_consolidated_format": 1, + "metadata": {".zgroup": GROUP, "a/.zarray": BAD}, +} +STORE_A = {".zarray": json.dumps(BAD).encode()} +STORE_G = {".zgroup": json.dumps(GROUP).encode()} +STORE_C = {".zmetadata": json.dumps(CONSOLIDATED).encode()} + + +def _accepts(read: Callable[..., object]) -> Callable[[Context | None], bool]: + """Whether `read`, given `context`, finds nothing wrong: True for a model, a reading with a model, an empty problem tuple, a guard saying yes, or a field read or left unclaimed.""" + + def accepted(context: Context | None) -> bool: + try: + found = read(context=context) + except zm.MetadataValidationError: + return False + if isinstance(found, tuple) and len(found) == 2 and isinstance(found[1], tuple): + return len(found[1]) == 0 # a resolved field and its problems + if isinstance(found, tuple): + return len(found) == 0 + if isinstance(found, bool): + return found + reading = getattr(found, "reading", found) + metadata = getattr(reading, "metadata", found) + return metadata is not None + + return accepted + + +ENTRY_POINTS: dict[str, Callable[..., object]] = { + "validate_array_metadata_v2": lambda context: zm.validate_array_metadata_v2( + BAD, context=context + ), + "is_array_metadata_v2": lambda context: zm.is_array_metadata_v2( + zm.parse_array_metadata_v2(BAD, context=EMPTY), context=context + ), + "parse_array_metadata_v2": lambda context: zm.parse_array_metadata_v2(BAD, context=context), + "read_array_metadata_v2": lambda context: zm.read_array_metadata_v2(BAD, context=context), + "ZarrV2ArrayMetadata": lambda context: zm.ZarrV2ArrayMetadata(BAD, context=context), + "ZarrV2ArrayMetadata.from_json": lambda context: zm.ZarrV2ArrayMetadata.from_json( + BAD, context=context + ), + "ZarrV2ArrayMetadata.from_key_value": lambda context: zm.ZarrV2ArrayMetadata.from_key_value( + STORE_A, context=context + ), + "ZarrV2ArrayMetadata.create_default": lambda context: zm.ZarrV2ArrayMetadata.create_default( + shape=(4,), chunks=(2,), compressor=BAD["compressor"], context=context + ), + "ZarrV2ArrayMetadata.with_context": lambda context: zm.ZarrV2ArrayMetadata( + BAD, context=EMPTY + ).with_context(context), + "ZarrV2ArrayMetadata.refined_in": lambda context: zm.ZarrV2ArrayMetadata( + BAD, context=EMPTY + ).refined_in(context), + "validate_group_metadata_v2": lambda context: zm.validate_group_metadata_v2( + GROUP, context=context + ), + "is_group_metadata_v2": lambda context: zm.is_group_metadata_v2(GROUP, context=context), + "parse_group_metadata_v2": lambda context: zm.parse_group_metadata_v2(GROUP, context=context), + "ZarrV2GroupMetadata": lambda context: zm.ZarrV2GroupMetadata(GROUP, context=context), + "ZarrV2GroupMetadata.from_json": lambda context: zm.ZarrV2GroupMetadata.from_json( + GROUP, context=context + ), + "ZarrV2GroupMetadata.from_key_value": lambda context: zm.ZarrV2GroupMetadata.from_key_value( + STORE_G, context=context + ), + "ZarrV2GroupMetadata.create_default": lambda context: zm.ZarrV2GroupMetadata.create_default( + context=context + ), + "ZarrV2ConsolidatedMetadata": lambda context: zm.ZarrV2ConsolidatedMetadata( + CONSOLIDATED, context=context + ), + "ZarrV2ConsolidatedMetadata.from_json": lambda context: zm.ZarrV2ConsolidatedMetadata.from_json( + CONSOLIDATED, context=context + ), + "ZarrV2ConsolidatedMetadata.from_key_value": ( + lambda context: zm.ZarrV2ConsolidatedMetadata.from_key_value(STORE_C, context=context) + ), + "read_repaired_consolidated_metadata_v2": ( + lambda context: zm.read_repaired_consolidated_metadata_v2(CONSOLIDATED, context=context) + ), + "resolve_dtype_v2": lambda context: resolve_dtype_v2(" None: + """Each v2 entry point accepts its document in an empty scope, and refuses the gzip level that is no integer in `CORE_V2`, which `None` names too, when the document holds it: the scope it is given is the scope it reads in.""" + accepted = _accepts(ENTRY_POINTS[name]) + assert accepted(EMPTY) is True + assert accepted(CORE_V2) is (name not in REFUSES_BAD) + assert accepted(None) is (name not in REFUSES_BAD) + + +@pytest.mark.parametrize("name", ENTRY_POINTS.keys()) +def test_error_every_v2_entry_point_refuses_a_scope_of_another_format(name: str) -> None: + """A v2 entry point given a v3 scope raises `TypeError`: a scope reads documents of one format, and a v3 scope claims nothing a v2 document writes.""" + with pytest.raises(TypeError, match="format"): + ENTRY_POINTS[name](context=CORE_AND_EXTENSIONS) diff --git a/packages/zarr-metadata/tests/model/test_sentinel.py b/packages/zarr-metadata/tests/model/test_sentinel.py index 252a4cdd30..1ec219d5b6 100644 --- a/packages/zarr-metadata/tests/model/test_sentinel.py +++ b/packages/zarr-metadata/tests/model/test_sentinel.py @@ -6,8 +6,8 @@ across process boundaries; these tests pin that behavior, since models hold `UNSET` as field values and must survive pickling and deep-copying. -The model round-trip tests compare whole structures: dataclass equality -compares every field, and `UNSET` compares by identity, so an impostor +The model round-trip tests compare whole structures: a model's equality +compares every member, and `UNSET` compares by identity, so an impostor sentinel produced by state-based pickling would fail the equality check. """ @@ -24,7 +24,6 @@ ZarrV2ArrayMetadata, ZarrV2GroupMetadata, ZarrV3ArrayMetadata, - ZarrV3ConsolidatedMetadata, ZarrV3GroupMetadata, ) @@ -34,8 +33,8 @@ # stay distinct from absent), and UNSET nested inside consolidated metadata. MODEL_CASES = { "array-v3-dimension-names-unset": ZarrV3ArrayMetadata.create_default(shape=(4,)), - "array-v3-dimension-names-set": ZarrV3ArrayMetadata.create_default(shape=(2, 2)).update( - dimension_names=("x", None) + "array-v3-dimension-names-set": ZarrV3ArrayMetadata.create_default( + shape=(2, 2), dimension_names=("x", None) ), "array-v2-attributes-unset": ZarrV2ArrayMetadata.create_default(shape=(4,)), "array-v2-attributes-empty": ZarrV2ArrayMetadata.create_default(shape=(4,), attributes={}), @@ -43,12 +42,14 @@ "group-v2-attributes-set": ZarrV2GroupMetadata.create_default(attributes={"a": 1}), "group-v3-consolidated-unset": ZarrV3GroupMetadata.create_default(), "group-v3-consolidated-with-unset-inside": ZarrV3GroupMetadata.create_default( - consolidated_metadata=ZarrV3ConsolidatedMetadata( - metadata={ - "child": ZarrV3ArrayMetadata.create_default(shape=(4,)), - "subgroup": ZarrV3GroupMetadata.create_default(), - } - ) + consolidated_metadata={ + "kind": "inline", + "must_understand": False, + "metadata": { + "child": ZarrV3ArrayMetadata.create_default(shape=(4,)).to_json(), + "subgroup": ZarrV3GroupMetadata.create_default().to_json(), + }, + }, ), } diff --git a/packages/zarr-metadata/tests/model/test_store_json.py b/packages/zarr-metadata/tests/model/test_store_json.py index 76bfd252b4..3192254fb2 100644 --- a/packages/zarr-metadata/tests/model/test_store_json.py +++ b/packages/zarr-metadata/tests/model/test_store_json.py @@ -4,7 +4,7 @@ so an attribute may hold `NaN`, `Infinity` or `-Infinity`. The model reads such a store, validates it, and writes it back the same way. Wherever the spec interprets a value, a non-finite number is refused when read, and a -document the reader would refuse is not written. +document the reader would refuse makes no model, so it is not written. """ from __future__ import annotations @@ -295,34 +295,44 @@ def test_error_a_non_finite_number_outside_attributes_is_located_when_read( @pytest.mark.parametrize( - ("model", "problems"), + ("build", "problems"), [ ( - ZarrV3ArrayMetadata.create_default(fill_value=math.nan, attributes={"x": math.nan}), + lambda: ZarrV3ArrayMetadata.create_default( + fill_value=math.nan, attributes={"x": math.nan} + ), [(("fill_value",), "invalid_value")], ), ( - ZarrV3GroupMetadata.create_default(attributes={"s": {1, 2}}), # pyright: ignore[reportArgumentType] + lambda: ZarrV3GroupMetadata.create_default(attributes={"s": {1, 2}}), # pyright: ignore[reportArgumentType] [(("attributes", "s"), "invalid_type")], ), ( - ZarrV3GroupMetadata.create_default(attributes={1: "a"}), # pyright: ignore[reportArgumentType] + lambda: ZarrV3GroupMetadata.create_default(attributes={1: "a"}), # pyright: ignore[reportArgumentType] [(("attributes",), "invalid_type")], ), ( - ZarrV3ArrayMetadata.create_default(shape=(2,), dimension_names=("x", "y")), + lambda: ZarrV3ArrayMetadata.create_default(shape=(2,), dimension_names=("x", "y")), [(("dimension_names",), "invalid_value")], ), ], ids=["non-finite-fill-value", "not-json", "non-string-key", "dimension-names-past-shape"], ) def test_error_a_document_the_reader_refuses_is_not_written( - model: _Stored, problems: list[tuple[tuple[str | int, ...], str]] + build: Callable[[], _Stored], problems: list[tuple[tuple[str | int, ...], str]] ) -> None: - # A model built by hand is not validated; the writer validates what it - # writes as the reader does. A non-string key was written as a string, - # a value that is not JSON raised `TypeError`, and the last was written - # and then refused on read. + # A model comes from a read of its document, so one the reader refuses + # makes no model to write. Written, a non-string key became a string, a + # value that is not JSON raised `TypeError`, and the last was refused + # only when read again. with pytest.raises(MetadataValidationError) as raised: - model.to_key_value() + build() assert [(problem.loc, problem.kind) for problem in raised.value.problems] == problems + + +def test_error_store_bytes_nested_deeper_than_python_reads_are_invalid_json() -> None: + # `json.loads` gives up on a hundred thousand `[` with a `RecursionError`, + # which is an ingestion failure like any other. + with pytest.raises(MetadataValidationError) as raised: + ZarrV3ArrayMetadata.from_key_value({"zarr.json": b"[" * 100_000}) + assert [(p.loc, p.kind) for p in raised.value.problems] == [(("zarr.json",), "invalid_json")] diff --git a/packages/zarr-metadata/tests/model/test_v2_scope_reads.py b/packages/zarr-metadata/tests/model/test_v2_scope_reads.py new file mode 100644 index 0000000000..03a2975d0f --- /dev/null +++ b/packages/zarr-metadata/tests/model/test_v2_scope_reads.py @@ -0,0 +1,106 @@ +"""The v2 array validator reads its dtype, codecs and fill value in the v2 scope.""" + +from __future__ import annotations + +from typing import Any + +import pytest + +from zarr_metadata.model import ( + MetadataValidationError, + ZarrV2ArrayMetadata, + validate_array_metadata_v2, +) + +Loc = tuple[str | int, ...] +BASE: dict[str, Any] = dict(ZarrV2ArrayMetadata.create_default().to_json()) + + +@pytest.mark.parametrize( + "changes", + [ + {"dtype": " None: + """A dtype the scope reads with a fill value its family takes, a compressor and filters the scope reads or leaves unclaimed, and a null fill value, each validate with no problem.""" + assert validate_array_metadata_v2({**BASE, **changes}) == () + + +@pytest.mark.parametrize( + ("changes", "at", "kind"), + [ + ({"dtype": "float32"}, ("dtype",), "invalid_value"), + ({"dtype": " None: + """A dtype, fill value, compressor or filter the scope refuses is one problem of the document, at the field or the parameter that is wrong; before, the string content of a dtype and the parameters of a codec went unjudged.""" + problems = validate_array_metadata_v2({**BASE, **changes}) + assert [(p.loc, p.kind) for p in problems] == [(at, kind)] + + +def test_create_default_keeps_zero_where_the_family_takes_it() -> None: + """`create_default` given a dtype and no fill value keeps `0` when the family takes it, and takes `null` otherwise; a fill value given is kept.""" + assert ZarrV2ArrayMetadata.create_default(dtype="|b1").fill_value is None + assert ZarrV2ArrayMetadata.create_default(dtype=" None: + """`ZarrV2ArrayMetadata` checks itself when built, so a dtype the scope refuses raises as any other problem does.""" + with pytest.raises(MetadataValidationError, match="typestr"): + ZarrV2ArrayMetadata.create_default(dtype="float32") diff --git a/packages/zarr-metadata/tests/test_json_schema.py b/packages/zarr-metadata/tests/test_json_schema.py new file mode 100644 index 0000000000..bf72501d32 --- /dev/null +++ b/packages/zarr-metadata/tests/test_json_schema.py @@ -0,0 +1,1015 @@ +"""JSON Schemas of what the package reads: a TypedDict, a field in a scope, a `zarr.json`. + +A schema is held to the reader it is written from: what the reader finds +nothing wrong with, the schema accepts, and what the type says -- a +bound, a key a closed TypedDict does not declare, a name a scope claims +-- the schema says too. What only a rule says, the schema does not. +""" + +from __future__ import annotations + +import json +import types +from collections.abc import Callable, Mapping +from typing import Annotated, Any, Literal, NewType, NotRequired, cast + +import pytest +from annotated_types import Ge, Gt, Interval, Le, Lt, MinLen +from hypothesis import HealthCheck, event, given, settings +from hypothesis import strategies as st +from jsonschema import Draft202012Validator +from typing_extensions import Doc, TypeAliasType, TypedDict + +from tests.v3.test_every_definition import CASES, KINDS +from zarr_metadata._common import JSONValue +from zarr_metadata._typed_json import Schemas +from zarr_metadata.model import ( + node_metadata_json_schema_v3, + validate_node_metadata_v3, +) +from zarr_metadata.typed_json import ( + check, + json_schema, +) +from zarr_metadata.v3._definition import ( + field_json_schema, +) +from zarr_metadata.v3.codec.gzip import GzipCodecConfiguration +from zarr_metadata.v3.definition import ( + CORE, + CORE_AND_EXTENSIONS, + ChunkGridDefinition, + ChunkKeyEncodingDefinition, + CodecDefinition, + Context, + DataTypeDefinition, + Definition, + StorageTransformerDefinition, + resolve, +) + +DIALECT = "https://json-schema.org/draft/2020-12/schema" + +# Built at run time, where a type checker reads each special form's +# arguments as it would in an annotation. +_annotated: Any = Annotated +_new_type: Callable[[str, object], object] = cast("Any", NewType) +_alias_type: Callable[[str, object], object] = cast("Any", TypeAliasType) + +Level = TypeAliasType("Level", Annotated[int, Interval(ge=0, le=9)]) +Tree = TypeAliasType("Tree", "int | tuple[Tree, ...]") +Name = NewType("Name", str) +Count = _new_type("Count", Annotated[int, Ge(0)]) + + +class Inner(TypedDict, closed=True): + x: int + + +class Open(TypedDict, closed=False): + x: NotRequired[int] + + +class Extra(TypedDict, extra_items=int): + x: str + + +class ExtraJSON(TypedDict, extra_items=JSONValue): + x: str + + +class Node(TypedDict, closed=True): + children: tuple[Node, ...] + + +def _closed(name: str, annotations: dict[str, object]) -> type: + """A closed TypedDict of these required keys, made here, so its annotations resolve in this module.""" + base: object = TypedDict + namespace = {"__annotations__": annotations, "__module__": __name__} + return types.new_class(name, (base,), {"closed": True}, lambda body: body.update(namespace)) + + +def _holding(annotation: object) -> type: + """A closed TypedDict of one required key, `value`, holding `annotation`.""" + return _closed("Holder", {"value": annotation}) + + +# Another class of the name `Inner`, which the schema tells from the first. +Both = _closed("Both", {"first": Inner, "second": _closed("Inner", {"y": str})}) + + +def _held(member: dict[str, Any], defs: dict[str, Any] | None = None) -> dict[str, Any]: + """The schema of `_holding` an annotation whose schema is `member`.""" + schema: dict[str, Any] = { + "$schema": DIALECT, + "type": "object", + "properties": {"value": member}, + "required": ["value"], + "additionalProperties": False, + } + return schema if defs is None else {**schema, "$defs": defs} + + +INNER = { + "type": "object", + "properties": {"x": {"type": "integer"}}, + "required": ["x"], + "additionalProperties": False, +} + +SHAPES: list[tuple[str, type, dict[str, Any]]] = [ + ("int", _holding(int), _held({"type": "integer"})), + ("float", _holding(float), _held({"type": "number"})), + ("bool", _holding(bool), _held({"type": "boolean"})), + ("str", _holding(str), _held({"type": "string"})), + ("null", _holding(None), _held({"type": "null"})), + ("json", _holding(JSONValue), _held({})), + ("one-value", _holding(Literal["a"]), _held({"const": "a"})), + # Sorted as the checker sorts them, so the order written does not show. + ("values", _holding(Literal["b", "a", 1]), _held({"enum": ["a", "b", 1]})), + ("array", _holding(tuple[int, ...]), _held({"type": "array", "items": {"type": "integer"}})), + ("array-of-json", _holding(tuple[JSONValue, ...]), _held({"type": "array"})), + ( + "pair", + _holding(tuple[int, str]), + _held( + { + "type": "array", + "prefixItems": [{"type": "integer"}, {"type": "string"}], + "items": False, + "minItems": 2, + } + ), + ), + ("empty-array", _holding(tuple[()]), _held({"type": "array", "maxItems": 0})), + ( + "union", + _holding(int | None), + _held({"anyOf": [{"type": "integer"}, {"type": "null"}]}), + ), + ( + "mapping", + _holding(Mapping[str, int]), + _held({"type": "object", "additionalProperties": {"type": "integer"}}), + ), + ("mapping-of-json", _holding(Mapping[str, JSONValue]), _held({"type": "object"})), + ("newtype", _holding(Name), _held({"type": "string"})), + ( + "interval", + _holding(Annotated[int, Interval(ge=0, le=9)]), + _held({"type": "integer", "minimum": 0, "maximum": 9}), + ), + ( + "exclusive-bounds", + _holding(Annotated[float, Gt(0), Lt(1)]), + _held({"type": "number", "exclusiveMinimum": 0, "exclusiveMaximum": 1}), + ), + ( + "doc", + _holding(Annotated[int, Ge(0), Doc("a count")]), + _held({"type": "integer", "description": "a count", "minimum": 0}), + ), + # A note that is no `Doc` says nothing a schema writes. + ("note", _holding(Annotated[str, "a note"]), _held({"type": "string"})), + # A value is held to its type's bounds and to its `NewType`'s: of two + # of one keyword, the stricter. + ( + "bounds-on-a-bounded-newtype", + _holding(_annotated[Count, Ge(-5), Lt(9)]), + _held({"type": "integer", "minimum": 0, "exclusiveMaximum": 9}), + ), + ( + "alias", + _holding(Level), + _held( + {"$ref": "#/$defs/Level"}, + {"Level": {"type": "integer", "minimum": 0, "maximum": 9}}, + ), + ), + ( + "alias-holding-itself", + _holding(Tree), + _held( + {"$ref": "#/$defs/Tree"}, + { + "Tree": { + "anyOf": [ + {"type": "integer"}, + {"type": "array", "items": {"$ref": "#/$defs/Tree"}}, + ] + } + }, + ), + ), + ("typeddict", _holding(Inner), _held({"$ref": "#/$defs/Inner"}, {"Inner": INNER})), + ( + "open", + Open, + {"$schema": DIALECT, "type": "object", "properties": {"x": {"type": "integer"}}}, + ), + ( + "extra-items", + Extra, + { + "$schema": DIALECT, + "type": "object", + "properties": {"x": {"type": "string"}}, + "required": ["x"], + "additionalProperties": {"type": "integer"}, + }, + ), + ( + "extra-items-of-json", + ExtraJSON, + { + "$schema": DIALECT, + "type": "object", + "properties": {"x": {"type": "string"}}, + "required": ["x"], + }, + ), + # A TypedDict that holds itself is referred to, not written in place. + ( + "holding-itself", + Node, + { + "$schema": DIALECT, + "$ref": "#/$defs/Node", + "$defs": { + "Node": { + "type": "object", + "properties": { + "children": {"type": "array", "items": {"$ref": "#/$defs/Node"}} + }, + "required": ["children"], + "additionalProperties": False, + } + }, + }, + ), + ( + "one-name-two-classes", + Both, + { + "$schema": DIALECT, + "type": "object", + "properties": { + "first": {"$ref": "#/$defs/Inner"}, + "second": {"$ref": "#/$defs/Inner2"}, + }, + "required": ["first", "second"], + "additionalProperties": False, + "$defs": { + "Inner": INNER, + "Inner2": { + "type": "object", + "properties": {"y": {"type": "string"}}, + "required": ["y"], + "additionalProperties": False, + }, + }, + }, + ), + ( + "a-definition-s-configuration", + GzipCodecConfiguration, + { + "$schema": DIALECT, + "type": "object", + "properties": {"level": {"type": "integer", "minimum": 0, "maximum": 9}}, + "required": ["level"], + "additionalProperties": False, + }, + ), +] + + +@pytest.mark.parametrize( + ("shape", "expected"), [case[1:] for case in SHAPES], ids=[case[0] for case in SHAPES] +) +def test_json_schema_writes_what_check_reads(shape: type, expected: dict[str, Any]) -> None: + schema = json_schema(shape) + assert schema == expected + Draft202012Validator.check_schema(schema) + assert json.loads(json.dumps(schema)) == schema + + +def test_json_schema_refuses_what_is_not_a_typeddict() -> None: + with pytest.raises(TypeError, match="is not a TypedDict"): + json_schema(int) + + +def test_json_schema_refuses_what_check_cannot_read() -> None: + unread = _holding(bytes) + with pytest.raises(TypeError) as raised: + check({}, unread) + with pytest.raises(TypeError, match=str(raised.value)): + json_schema(unread) + + +def test_error_a_schema_that_fails_to_write_leaves_nothing_behind() -> None: + # Reserved before it is written, so that it can refer to itself, and + # given up when the writing fails, so that nothing refers to an empty + # schema, which takes anything. + schemas = Schemas() + unread = _closed("Unread", {"value": bytes}) + with pytest.raises(TypeError, match="is not a shape JSON takes"): + schemas.of(unread) + assert schemas.document({}) == {"$schema": DIALECT} + with pytest.raises(TypeError, match="is not a shape JSON takes"): + schemas.of(unread) + # Nor is what it wrote of a shape it holds before it failed, nor a use + # it counted: a later root of that shape, used once, is written in place. + held = _closed("Held", {"x": int}) + with pytest.raises(TypeError, match="is not a shape JSON takes"): + schemas.of(_closed("Partly", {"held": held, "value": bytes})) + assert schemas.document({}) == {"$schema": DIALECT} + assert schemas.document(schemas.of(held)) == { + "$schema": DIALECT, + **Schemas().object_of(held), + } + + +def test_json_schema_refuses_metadata_check_does_not_hold_a_value_to() -> None: + with pytest.raises(TypeError, match="is not a constraint the checker reads"): + json_schema(_holding(Annotated[tuple[int, ...], MinLen(1)])) + + +_PROPERTY = settings(max_examples=300, deadline=None, suppress_health_check=[HealthCheck.too_slow]) + +_MARKERS: dict[str, Callable[[int], object]] = {"ge": Ge, "gt": Gt, "le": Le, "lt": Lt} +_LOW = st.none() | st.tuples(st.sampled_from(("ge", "gt")), st.integers(-3, 3)) +_HIGH = st.none() | st.tuples(st.sampled_from(("le", "lt")), st.integers(-3, 3)) +_LAYERS = st.lists( + st.tuples(st.sampled_from(("annotated", "newtype", "alias")), _LOW, _HIGH), + min_size=1, + max_size=3, +) +_NUMBERS = st.integers(-5, 5) | st.floats(-5, 5, allow_nan=False) | st.sampled_from((True, "1")) + + +def _integral_as_int(value: object) -> object: + """`value` as JSON Schema reads a number: `1.0` is the integer 1.""" + return int(value) if isinstance(value, float) and value.is_integer() else value + + +@_PROPERTY +@given(base=st.sampled_from((int, float)), layers=_LAYERS, values=st.lists(_NUMBERS, max_size=8)) +def test_bounds_in_layers_are_written_as_check_holds_a_value_to_them( + base: type, layers: list[tuple[str, object, object]], values: list[object] +) -> None: + # A number's bounds, on it, on a `NewType` of it, on an alias of it, + # layer on layer: the schema accepts what `check` does, a number with no + # fraction read as the integer it equals. + annotation: object = base + for index, (how, low, high) in enumerate(layers): + sides = [cast("tuple[str, int]", side) for side in (low, high) if side is not None] + bounds = [_MARKERS[name](bound) for name, bound in sides] + bounded = _annotated[(annotation, *bounds)] if len(bounds) != 0 else annotation + if how == "newtype": + annotation = _new_type(f"Bounded{index}", bounded) + elif how == "alias": + annotation = _alias_type(f"Bounded{index}", bounded) + else: + annotation = bounded + holder = _holding(annotation) + try: + schema = json_schema(holder) + except TypeError: + # A second bound from one side, which neither reads. + event("refused") + with pytest.raises(TypeError): + check({}, holder) + return + validator = Draft202012Validator(schema) + for value in values: + accepted = check({"value": _integral_as_int(value)}, holder)[1] == () + event("accepted" if accepted else "refused a value") + assert validator.is_valid(cast("Any", {"value": value})) == accepted, (value, schema) + + +# --- values changed in one place -------------------------------------------- + +_GONE = object() +_NAMES = ("r16", "r16\n", "r*", "gzip", "bytes", "crc32c", "int8", "regular", "default", "acme.x") +_KEYS = ("name", "configuration", "must_understand", "level", "x") +_SCALARS = ( + st.none() + | st.booleans() + | st.integers(-3, 300) + | st.floats(-3, 3, allow_nan=False) + | st.sampled_from(_NAMES) + | st.text(max_size=3) +) +_VALUES = st.recursive( + _SCALARS, + lambda inner: ( + st.lists(inner, max_size=2) + | st.dictionaries(st.sampled_from(_KEYS) | st.text(max_size=2), inner, max_size=2) + ), + max_leaves=4, +) + + +def _places(value: object, path: tuple[str | int, ...] = ()) -> list[tuple[str | int, ...]]: + """Every place in `value`: itself, and each member or entry inside it, depth first.""" + found = [path] + if isinstance(value, dict): + for key, entry in cast("dict[str, object]", value).items(): + found += _places(entry, (*path, key)) + elif isinstance(value, list): + for index, entry in enumerate(cast("list[object]", value)): + found += _places(entry, (*path, index)) + return found + + +def _at(value: object, path: tuple[str | int, ...]) -> object: + for step in path: + value = cast("Any", value)[step] + return value + + +def _put(value: object, path: tuple[str | int, ...], new: object) -> object: + """`value` with what sits at `path` replaced by `new`, or removed when `new` is `_GONE`.""" + if len(path) == 0: + return new + step, rest = path[0], path[1:] + copy: Any = ( + dict(cast("dict[str, object]", value)) + if isinstance(value, dict) + else list(cast("list[object]", value)) + ) + if len(rest) == 0 and new is _GONE: + del copy[step] + else: + copy[step] = _put(copy[step], rest, new) + return copy + + +@st.composite +def _changed(draw: st.DrawFn, value: object) -> object: + """`value` changed in one or two places: something replaced, dropped or added, a string with a character more, or a field renamed.""" + for _ in range(draw(st.integers(1, 2))): + path = draw(st.sampled_from(_places(value))) + here = _at(value, path) + how = draw(st.sampled_from(("replace", "drop", "add", "extend", "rename"))) + if how == "drop" and len(path) != 0: + value = _put(value, path, _GONE) + elif how == "add" and isinstance(here, dict): + members = cast("dict[str, object]", here) + value = _put(value, path, {**members, draw(st.sampled_from(_KEYS)): draw(_VALUES)}) + elif how == "add" and isinstance(here, list): + value = _put(value, path, [*cast("list[object]", here), draw(_VALUES)]) + elif how == "extend" and isinstance(here, str): + value = _put(value, path, here + draw(st.sampled_from(("\n", " ", "0", "x")))) + elif how == "rename" and isinstance(here, dict) and "name" in here: + value = _put(value, (*path, "name"), draw(st.sampled_from(_NAMES))) + elif how == "rename" and isinstance(here, str): + value = _put(value, path, draw(st.sampled_from(_NAMES))) + else: + value = _put(value, path, draw(_VALUES)) + return value + + +_SCOPES = (CORE, CORE_AND_EXTENSIONS, Context.of()) +_FIELD_BASES: tuple[tuple[str, object], ...] = ( + *CASES, + # Names nothing claims, whose configuration nothing judges. + ("codecs:acme.x", {"name": "acme.x", "configuration": {"x": [1]}}), + ("data_type:acme.t", {"name": "acme.t", "configuration": {"bits": 8}}), + ("chunk_grid:acme.g", {"name": "acme.g", "configuration": {"chunk_shape": [0]}}), +) +_FIELD_SCHEMAS: dict[tuple[type[Definition[Any]], int], Draft202012Validator] = {} + + +def _field_validator(kind: type[Definition[Any]], scope: int) -> Draft202012Validator: + if (kind, scope) not in _FIELD_SCHEMAS: + schema = field_json_schema(kind, _SCOPES[scope]) + _FIELD_SCHEMAS[(kind, scope)] = Draft202012Validator(schema) + return _FIELD_SCHEMAS[(kind, scope)] + + +# --- a field in a scope ----------------------------------------------------- + +_INDEXED = {"chunk_shape": [2], "codecs": ["bytes"]} + +# Each field, the kind and scope it is read in, whether the schema accepts +# it, and whether the scope reads it without a problem. Where the two +# differ, a rule found what the schema cannot say. +FIELDS: list[tuple[str, object, type[Definition[Any]], Context, bool, bool]] = [ + ("object", {"name": "gzip", "configuration": {"level": 5}}, CodecDefinition, CORE, True, True), + ( + "out-of-bounds", + {"name": "gzip", "configuration": {"level": 12}}, + CodecDefinition, + CORE, + False, + False, + ), + ( + "unknown-key", + {"name": "gzip", "configuration": {"level": 5, "window": 15}}, + CodecDefinition, + CORE, + False, + False, + ), + ("bare-name-needing-a-configuration", "gzip", CodecDefinition, CORE, False, False), + ("bare-name", "bytes", CodecDefinition, CORE, True, True), + ( + "understood", + {"name": "gzip", "configuration": {"level": 5}, "must_understand": True}, + CodecDefinition, + CORE, + True, + True, + ), + ( + "not-understood", + {"name": "gzip", "configuration": {"level": 5}, "must_understand": False}, + CodecDefinition, + CORE, + False, + False, + ), + ( + "stray-member", + {"name": "gzip", "configuration": {"level": 5}, "version": 2}, + CodecDefinition, + CORE, + False, + False, + ), + ( + "unclaimed", + {"name": "zfpy", "configuration": {"mode": 4}}, + CodecDefinition, + CORE, + True, + True, + ), + ("unclaimed-bare", "zfpy", CodecDefinition, CORE, True, True), + ( + "unclaimed-not-understood", + {"name": "zfpy", "must_understand": False}, + CodecDefinition, + CORE, + False, + False, + ), + # Nothing in an empty scope claims gzip, so nothing judges it. + ( + "out-of-scope", + {"name": "gzip", "configuration": {"level": 12}}, + CodecDefinition, + Context.of(), + True, + True, + ), + ( + "a-rule", + { + "name": "blosc", + "configuration": {"cname": "lz4", "clevel": 1, "shuffle": "shuffle", "blocksize": 0}, + }, + CodecDefinition, + CORE, + True, + False, + ), + ( + "static-index-codec", + { + "name": "sharding_indexed", + "configuration": {**_INDEXED, "index_codecs": ["bytes", "crc32c"]}, + }, + CodecDefinition, + CORE, + True, + True, + ), + ( + "dynamic-index-codec", + { + "name": "sharding_indexed", + "configuration": { + **_INDEXED, + "index_codecs": [{"name": "gzip", "configuration": {"level": 1}}], + }, + }, + CodecDefinition, + CORE, + False, + False, + ), + ( + "unclaimed-index-codec", + { + "name": "sharding_indexed", + "configuration": {**_INDEXED, "index_codecs": ["bytes", "acme.sum"]}, + }, + CodecDefinition, + CORE, + True, + True, + ), + ( + "inner-codec-out-of-bounds", + { + "name": "sharding_indexed", + "configuration": { + **_INDEXED, + "codecs": ["bytes", {"name": "gzip", "configuration": {"level": 12}}], + "index_codecs": ["bytes"], + }, + }, + CodecDefinition, + CORE, + False, + False, + ), + ("raw-bits", "r16", DataTypeDefinition, CORE, True, True), + ("raw-bits-object", {"name": "r16", "configuration": {}}, DataTypeDefinition, CORE, True, True), + # What a raw-bits name carries is not written beside it. + ( + "raw-bits-configured", + {"name": "r16", "configuration": {"bits": 16}}, + DataTypeDefinition, + CORE, + False, + False, + ), + ("raw-bits-of-a-size-the-spec-refuses", "r12", DataTypeDefinition, CORE, True, False), + # `r*` is the table's notation, and no name at all: `*` is no character + # of one. + ("raw-bits-notation", "r*", DataTypeDefinition, CORE, False, False), + # Matched to the end of the name, as the package matches it, so a + # validator that matches as Python does takes no final newline for it; + # nor is it a name at all, a newline being no character of one. + ( + "raw-bits-and-a-newline", + {"name": "r16\n", "configuration": {"x": 1}}, + DataTypeDefinition, + CORE, + False, + False, + ), + ( + "struct-field-out-of-bounds", + { + "name": "struct", + "configuration": { + "fields": [ + { + "name": "acme.t", + "data_type": { + "name": "numpy.datetime64", + "configuration": {"unit": "s", "scale_factor": 0}, + }, + } + ] + }, + }, + DataTypeDefinition, + CORE_AND_EXTENSIONS, + False, + False, + ), + ( + "grid-out-of-bounds", + {"name": "regular", "configuration": {"chunk_shape": [0]}}, + ChunkGridDefinition, + CORE, + False, + False, + ), + ( + "encoding", + {"name": "v2", "configuration": {"separator": "/"}}, + ChunkKeyEncodingDefinition, + CORE, + True, + True, + ), + ("transformer", {"name": "acme.cache"}, StorageTransformerDefinition, CORE, True, True), +] + + +@pytest.mark.parametrize( + ("field", "kind", "scope", "accepted", "clean"), + [case[1:] for case in FIELDS], + ids=[case[0] for case in FIELDS], +) +def test_field_json_schema_says_what_the_scope_reads( + field: object, kind: type[Definition[Any]], scope: Context, accepted: bool, clean: bool +) -> None: + schema = field_json_schema(kind, scope) + Draft202012Validator.check_schema(schema) + assert Draft202012Validator(schema).is_valid(cast("Any", field)) == accepted + assert (resolve(field, kind, scope)[1] == ()) == clean + + +@pytest.mark.parametrize( + ("key", "field"), CASES, ids=[f"{key}:{index}" for index, (key, _) in enumerate(CASES)] +) +def test_every_example_field_is_one_its_schema_accepts(key: str, field: object) -> None: + schema = field_json_schema(KINDS[key.split(":")[0]], CORE_AND_EXTENSIONS) + assert Draft202012Validator(schema).is_valid(cast("Any", field)) + + +@_PROPERTY +@given(st.data()) +def test_a_field_read_without_a_problem_is_one_its_schema_accepts(data: st.DataObject) -> None: + # Each example changed in one place, in one of three scopes: whatever + # the scope still reads without a problem, the schema accepts. + key, example = data.draw(st.sampled_from(_FIELD_BASES), label="example") + kind = KINDS[key.split(":")[0]] + scope = data.draw(st.sampled_from(range(len(_SCOPES))), label="scope") + field = data.draw(_changed(json.loads(json.dumps(example))), label="field") + read = resolve(field, kind, _SCOPES[scope])[1] == () + event("read without a problem" if read else "a problem") + if read: + assert _field_validator(kind, scope).is_valid(cast("Any", field)) + + +def test_field_json_schema_refuses_what_is_not_a_kind() -> None: + with pytest.raises(TypeError, match="is not a kind of metadata"): + field_json_schema(Definition, CORE) + + +# --- a zarr.json ------------------------------------------------------------ + +ARRAY: dict[str, Any] = { + "zarr_format": 3, + "node_type": "array", + "shape": [4, 4], + "data_type": "int8", + "chunk_grid": {"name": "regular", "configuration": {"chunk_shape": [2, 2]}}, + "chunk_key_encoding": {"name": "default"}, + "fill_value": 0, + "codecs": [{"name": "bytes"}, {"name": "gzip", "configuration": {"level": 5}}], +} +FLOATS: dict[str, Any] = { + **ARRAY, + "data_type": "float32", + "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}], +} +GROUP: dict[str, Any] = {"zarr_format": 3, "node_type": "group"} + + +def _consolidated(**metadata: object) -> dict[str, Any]: + return { + **GROUP, + "consolidated_metadata": {"kind": "inline", "must_understand": False, "metadata": metadata}, + } + + +# Each document, whether the schema accepts it, and whether +# `validate_node_metadata_v3` finds nothing wrong with it. +DOCUMENTS: list[tuple[str, dict[str, Any], bool, bool]] = [ + ("array", ARRAY, True, True), + ("group", {**GROUP, "attributes": {"a": [1, None]}}, True, True), + ("fill-value-out-of-range", {**ARRAY, "fill_value": 300}, False, False), + ("fill-value-of-the-wrong-type", {**ARRAY, "fill_value": "0"}, False, False), + ("float-fill-value", {**FLOATS, "fill_value": "NaN"}, True, True), + ("hex-fill-value", {**FLOATS, "fill_value": "0x7fc00000"}, True, True), + # A string that is no float's is a rule's to find. + ("fill-value-a-rule-refuses", {**FLOATS, "fill_value": "nan"}, True, False), + ("unclaimed-data-type", {**ARRAY, "data_type": "acme.int7", "fill_value": "any"}, True, True), + ( + "a-name-raw-bits-and-a-newline", + {**ARRAY, "data_type": "r16\n", "fill_value": "any"}, + False, + False, + ), + ("raw-bits", {**ARRAY, "data_type": "r16", "fill_value": [0, 255]}, True, True), + ( + "raw-bits-byte-out-of-range", + {**ARRAY, "data_type": "r16", "fill_value": [0, 256]}, + False, + False, + ), + ( + "codec-not-understood", + {**ARRAY, "codecs": [{"name": "bytes", "must_understand": False}]}, + False, + False, + ), + ( + "codec-out-of-bounds", + {**ARRAY, "codecs": [{"name": "bytes"}, {"name": "gzip", "configuration": {"level": 12}}]}, + False, + False, + ), + ("negative-shape", {**ARRAY, "shape": [-1, 4]}, False, False), + ("wrong-format", {**ARRAY, "zarr_format": 2}, False, False), + ( + "no-node-type", + {key: value for key, value in ARRAY.items() if key != "node_type"}, + False, + False, + ), + ("an-extension-field", {**ARRAY, "acme": {"must_understand": False}}, True, True), + # What members read together say is the rules'. + ("names-for-another-rank", {**ARRAY, "dimension_names": ["x"]}, True, False), + ( + "codecs-out-of-order", + {**ARRAY, "codecs": [{"name": "gzip", "configuration": {"level": 5}}, "bytes"]}, + True, + False, + ), + ("consolidated", _consolidated(a=ARRAY, b=GROUP), True, True), + ("consolidated-null", {**GROUP, "consolidated_metadata": None}, False, False), + ("consolidated-bad-array", _consolidated(a={**ARRAY, "fill_value": 300}), False, False), + ( + "consolidated-within-consolidated", + _consolidated(b=_consolidated(c={**ARRAY, "fill_value": 300})), + False, + False, + ), + ( + "consolidated-kind", + { + **GROUP, + "consolidated_metadata": {"kind": "sidecar", "must_understand": False, "metadata": {}}, + }, + False, + False, + ), +] + +NODE_SCHEMA = node_metadata_json_schema_v3() + + +@pytest.mark.parametrize( + ("document", "accepted", "clean"), + [case[1:] for case in DOCUMENTS], + ids=[case[0] for case in DOCUMENTS], +) +def test_node_metadata_json_schema_says_what_a_zarr_json_holds( + document: dict[str, Any], accepted: bool, clean: bool +) -> None: + assert Draft202012Validator(NODE_SCHEMA).is_valid(document) == accepted + assert (validate_node_metadata_v3(document) == ()) == clean + + +_BASES: tuple[dict[str, Any], ...] = ( + ARRAY, + FLOATS, + {**ARRAY, "data_type": "r16", "fill_value": [0, 0]}, + { + **ARRAY, + "codecs": [ + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [1, 1], + "codecs": ["bytes", {"name": "gzip", "configuration": {"level": 1}}], + "index_codecs": ["bytes", "crc32c"], + }, + } + ], + }, + _consolidated(a=ARRAY, b=GROUP), +) +_NODE_SCHEMAS: dict[int, Draft202012Validator] = {} + + +def _node_validator(scope: int) -> Draft202012Validator: + if scope not in _NODE_SCHEMAS: + _NODE_SCHEMAS[scope] = Draft202012Validator( + node_metadata_json_schema_v3(context=_SCOPES[scope]) + ) + return _NODE_SCHEMAS[scope] + + +@_PROPERTY +@given(st.data()) +def test_a_document_read_without_a_problem_is_one_its_schema_accepts(data: st.DataObject) -> None: + # Each document changed in one place, in one of three scopes: whatever + # `validate_node_metadata_v3` finds nothing wrong with, the schema accepts. + base = data.draw(st.sampled_from(_BASES), label="base") + scope = data.draw(st.sampled_from(range(len(_SCOPES))), label="scope") + document = data.draw(_changed(json.loads(json.dumps(base))), label="document") + read = validate_node_metadata_v3(document, context=_SCOPES[scope]) == () + event("read without a problem" if read else "a problem") + if read: + assert _node_validator(scope).is_valid(cast("Any", document)) + + +@pytest.mark.parametrize( + "scope", + [CORE, CORE_AND_EXTENSIONS, Context.of()], + ids=["core", "core-and-extensions", "empty"], +) +def test_every_schema_is_a_json_schema_written_the_same_each_time(scope: Context) -> None: + schema = node_metadata_json_schema_v3(context=scope) + Draft202012Validator.check_schema(schema) + assert json.loads(json.dumps(schema)) == schema + assert schema == node_metadata_json_schema_v3(context=scope) + defs = cast("dict[str, JSONValue]", schema["$defs"]) + assert {"ZarrV3ArrayMetadataJSON", "ZarrV3GroupMetadataJSON"} <= defs.keys() + for kind in ( + CodecDefinition, + DataTypeDefinition, + ChunkGridDefinition, + ChunkKeyEncodingDefinition, + StorageTransformerDefinition, + ): + Draft202012Validator.check_schema(field_json_schema(kind, scope)) + + +@pytest.mark.parametrize( + ("inner", "outer", "expected"), + [ + (Gt(1), Gt(3), {"type": "integer", "exclusiveMinimum": 3}), + (Gt(3), Gt(1), {"type": "integer", "exclusiveMinimum": 3}), + (Lt(9), Lt(5), {"type": "integer", "exclusiveMaximum": 5}), + (Lt(5), Lt(9), {"type": "integer", "exclusiveMaximum": 5}), + (Le(5), Le(9), {"type": "integer", "maximum": 5}), + (Le(9), Le(5), {"type": "integer", "maximum": 5}), + (Ge(5), Ge(9), {"type": "integer", "minimum": 9}), + (Ge(9), Ge(5), {"type": "integer", "minimum": 9}), + ], + ids=[ + "gt-looser-inside", + "gt-stricter-inside", + "lt", + "lt-stricter-inside", + "le", + "le-stricter-inside", + "ge", + "ge-stricter-inside", + ], +) +def test_the_stricter_of_two_bounds_of_one_keyword_is_kept( + inner: object, outer: object, expected: dict[str, object] +) -> None: + # A bounded NewType bounded again: each keyword keeps the bound that + # admits less, whichever layer carries it. + # `NewType` takes no `Annotated` for pyright; it does at runtime. + Bounded = cast("Callable[[str, object], type]", NewType)("Bounded", _annotated[int, inner]) + assert json_schema(_holding(_annotated[Bounded, outer])) == _held(expected) + + +def test_a_stray_member_beside_an_unclaimed_name_is_refused_by_the_schema() -> None: + # As the reader refuses it: the envelope of a name nothing claims holds + # nothing but the name and its configuration. + field = {"name": "zfpy", "version": 2} + assert [(p.loc, p.kind) for p in resolve(field, CodecDefinition, CORE)[1]] == [ + (("version",), "unknown_key") + ] + assert not Draft202012Validator(field_json_schema(CodecDefinition, CORE)).is_valid(field) + + +@pytest.mark.parametrize( + "name", + [ + "zstd", + "numcodecs.adler32", + "vlen-utf8", + "acme_x", + "r16", + "https://example.com/c", + "urn:acme:c", + "urn:acme:%C3%BC", + "", + " ", + "Int8", + "9x", + "int8 ", + "int8\n", + "foo/bar", + "-int8", + "a", + "r*", + "x:", + "a: b", + "urn:acme:c\n", + "urn:acme:ü", + "urn:a\x1c", + "urn:a", + ], +) +def test_the_schema_names_an_extension_as_the_reader_does(name: str) -> None: + # The spec's regex, or a URI: what `well_named` accepts, the schema's + # `pattern` accepts, and what it refuses, the schema refuses. + field = {"name": name} + _, problems = resolve(field, CodecDefinition, CORE) + accepted = len(problems) == 0 + assert ( + Draft202012Validator(field_json_schema(CodecDefinition, CORE)).is_valid(field) is accepted + ) + assert accepted is ( + name + in ( + "zstd", + "numcodecs.adler32", + "vlen-utf8", + "acme_x", + "https://example.com/c", + "urn:acme:c", + "urn:acme:%C3%BC", + ) + or name == "r16" + ) diff --git a/packages/zarr-metadata/tests/test_problem_data.py b/packages/zarr-metadata/tests/test_problem_data.py new file mode 100644 index 0000000000..1fb1a9f611 --- /dev/null +++ b/packages/zarr-metadata/tests/test_problem_data.py @@ -0,0 +1,408 @@ +"""What a problem carries besides its message: what was found, and what was expected. + +As pydantic's errors and zod's issues do, a problem carries its message's +data: `input`, what the value handed in holds at the problem's `loc`, and +`ctx`, what was expected there, where that is more than a type -- a +type's bounds, a closed set's values. Every reader fills `input`, so a +rule says only where a problem is. +""" + +from __future__ import annotations + +import copy +import dataclasses +import math +import pickle +from types import MappingProxyType +from typing import TYPE_CHECKING, Annotated, Any, Literal, cast + +import pytest +from annotated_types import Le +from typing_extensions import TypedDict + +from zarr_metadata._json import ( + arrays_to_tuples, + is_canonical_json, + validate_json, + value_at, + with_input, + within, +) +from zarr_metadata._sentinel import UNSET +from zarr_metadata.model import ( + MetadataValidationError, + ValidationProblem, + ZarrV2ConsolidatedMetadata, + ZarrV3ArrayMetadata, + parse_array_metadata_v3, + read_array_metadata_v3, + validate_array_metadata_v2, + validate_array_metadata_v3, + validate_group_metadata_v2, + validate_group_metadata_v3, + validate_metadata_field_v3, + validate_node_metadata_v3, + validate_node_name_v3, + validate_node_path_v3, +) +from zarr_metadata.typed_json import ( + check, +) +from zarr_metadata.v3.codec.gzip import GZIP_CODEC +from zarr_metadata.v3.definition import ( + CORE_AND_EXTENSIONS, + CodecDefinition, + DataTypeDefinition, + fill_value_problems, + resolve, +) + +if TYPE_CHECKING: + from collections.abc import Callable, Sequence + + from zarr_metadata._common import JSONValue + + +@pytest.mark.parametrize( + ("found", "ctx"), + [ + (UNSET, {}), + (12, {"ge": 0, "le": 9}), + (None, {"expected": ["C", "F"]}), + ], + ids=["nothing-found", "a-bound", "a-null-found"], +) +def test_a_problem_carries_what_was_found_and_what_was_expected( + found: JSONValue | UNSET, ctx: dict[str, JSONValue] +) -> None: + written = dict(ctx) + problem = ValidationProblem(("level",), "bad level", "invalid_value", input=found, ctx=written) + assert problem.input is found + # Held as a read-only view of a copy, arrays as tuples. + assert dict(problem.ctx) == arrays_to_tuples(ctx) + # What it says, and where, is the problem; what it carries is detail. + bare = ValidationProblem(("level",), "bad level", "invalid_value") + assert problem == bare + assert hash(problem) == hash(bare) + assert repr(problem) == repr(bare) + assert str(problem) == "level: bad level" + # A raised error is a finished report. + written["more"] = 1 + assert dict(problem.ctx) == arrays_to_tuples(ctx) + with pytest.raises(TypeError): + problem.ctx["ge"] = 1 # pyright: ignore[reportIndexIssue] + + +@pytest.mark.parametrize( + "ctx", + [["ge", 0], {1: 0}, {"ge": {0}}, {"ge": math.nan}], + ids=["an-array", "a-key-that-is-not-a-string", "a-set", "nan"], +) +def test_error_a_problem_refuses_a_ctx_that_is_not_an_object_of_json_values(ctx: object) -> None: + with pytest.raises(TypeError, match="ctx is an object of JSON values"): + ValidationProblem(("level",), "bad level", "invalid_value", ctx=ctx) # pyright: ignore[reportArgumentType] + + +def test_an_error_about_what_is_not_json_pickles() -> None: + # A problem holds no input that is not JSON, which might not pickle, + # so the error it is raised in crosses to another process as it is. + value = {"shape": [lambda: 4]} + problems = validate_json(value) + assert [(problem.loc, problem.input) for problem in problems] == [(("shape", 0), UNSET)] + again = pickle.loads(pickle.dumps(MetadataValidationError(problems))) + assert again.problems == problems + + +def test_a_problem_pickles_and_copies_with_what_it_carries() -> None: + problem = ValidationProblem(("shape",), "bad", "invalid_value", input=[12], ctx={"ge": 0}) + error = MetadataValidationError([problem]) + for again in ( + pickle.loads(pickle.dumps(problem)), + copy.copy(problem), + copy.deepcopy(problem), + pickle.loads(pickle.dumps(error)).problems[0], + ): + assert again == problem + assert again.input == [12] + assert dict(again.ctx) == {"ge": 0} + with pytest.raises(TypeError): + again.ctx["ge"] = 1 # pyright: ignore[reportIndexIssue] + + +def _deep(levels: int) -> list[object]: + """An array nested `levels` deep, deeper than the interpreter walks.""" + value: list[object] = [] + for _ in range(levels): + value = [value] + return value + + +@pytest.mark.parametrize( + ("value", "at", "loc", "held", "found"), + [ + ({"a": [1, {"b": 2}]}, (), ("a", 1, "b"), UNSET, 2), + ({"a": None}, (), ("a",), UNSET, None), + ({"a": 1}, (), ("b",), UNSET, UNSET), + ({"a": [1]}, (), ("a", 1), UNSET, UNSET), + ({"a": "text"}, (), ("a", 0), UNSET, UNSET), + ({"a": [1]}, (), ("a", "b"), UNSET, UNSET), + ([1, 2], ("codecs",), ("codecs", 1), UNSET, 2), + ([1, 2], ("codecs",), ("storage_transformers", 1), UNSET, UNSET), + ({"a": {1, 2}}, (), ("a",), UNSET, UNSET), + ({"a": _deep(100_000)}, (), ("a",), UNSET, UNSET), + ({"a": 1}, (), ("a",), "what a reader inside found", 1), + ({"a": {1, 2}}, (), ("a",), [1, 2], [1, 2]), + ], + ids=[ + "nested", + "a-null", + "a-missing-key", + "past-the-end", + "a-string-is-no-array", + "an-array-has-no-keys", + "under-where-the-value-sits", + "not-under-it", + "what-is-not-json", + "too-deep-to-walk", + "the-value-handed-in-wins", + "what-is-not-json-keeps-what-it-holds", + ], +) +def test_a_problem_holds_what_its_value_holds_at_its_loc( + value: object, + at: tuple[str | int, ...], + loc: tuple[str | int, ...], + held: JSONValue | UNSET, + found: JSONValue | UNSET, +) -> None: + (problem,) = with_input([ValidationProblem(loc, "bad", "invalid_value", input=held)], value, at) + assert problem.input == found + assert (problem.input is UNSET) is (found is UNSET) + + +class GzipLevelOnly(TypedDict, closed=True): + level: int + + +def _raised(read: Callable[[], object]) -> Sequence[ValidationProblem]: + with pytest.raises(MetadataValidationError) as caught: + read() + return caught.value.problems + + +BAD_ARRAY: dict[str, object] = { + "zarr_format": 3, + "node_type": "array", + "shape": [4, 4], + "data_type": "uint8", + "chunk_grid": {"name": "regular", "configuration": {"chunk_shape": [2]}}, + "chunk_key_encoding": {"name": "default"}, + "fill_value": 300, + "codecs": [ + {"name": "transpose", "configuration": {"order": [0, 1, 2]}}, + {"name": "bytes"}, + {"name": "gzip", "configuration": {"level": 12, "extra": [1]}}, + ], + "dimension_names": ["x"], + "attributes": {}, +} +"""An array document as `json.loads` gives one, arrays as lists, with a problem in each part a reader reads.""" + +BAD_ARRAY_LOCS = { + ("chunk_grid", "configuration", "chunk_shape"), + ("fill_value",), + ("codecs", 0, "configuration", "order"), + ("codecs", 2, "configuration", "level"), + ("codecs", 2, "configuration", "extra"), + ("dimension_names",), +} + +READERS: list[tuple[Callable[[object], Sequence[ValidationProblem]], object]] = [ + (lambda value: check(value, GzipLevelOnly)[1], {"level": 12, "extra": [1]}), + (lambda value: GZIP_CODEC.read_configuration(value)[1], {"level": 12, "extra": [1]}), + ( + lambda value: resolve(value, CodecDefinition, CORE_AND_EXTENSIONS)[1], + {"name": "gzip", "configuration": {"level": 12}, "must_understand": "yes"}, + ), + ( + lambda value: fill_value_problems( + resolve("uint8", DataTypeDefinition, CORE_AND_EXTENSIONS)[0], value + ), + 300, + ), + (validate_json, {"a": [math.inf]}), + ( + validate_metadata_field_v3, + {"name": 1, "configuration": {"a": [math.nan]}, "extra": [1]}, + ), + (validate_node_name_v3, "__/"), + (validate_node_path_v3, "a//"), + (validate_array_metadata_v3, BAD_ARRAY), + (lambda value: read_array_metadata_v3(value).problems, BAD_ARRAY), + (lambda value: _raised(lambda: parse_array_metadata_v3(value)), BAD_ARRAY), + (lambda value: _raised(lambda: ZarrV3ArrayMetadata.from_json(value)), BAD_ARRAY), + (validate_node_metadata_v3, BAD_ARRAY), + ( + validate_group_metadata_v3, + { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": { + "kind": "inline", + "must_understand": [False], + "metadata": {"a": BAD_ARRAY, "__b": BAD_ARRAY, "a/c": BAD_ARRAY, "d/e": BAD_ARRAY}, + }, + }, + ), + ( + validate_array_metadata_v2, + {"zarr_format": 2, "shape": [4], "chunks": [2, 2], "order": "Q", "filters": [[1]]}, + ), + (validate_group_metadata_v2, {"zarr_format": 3, "extra": [1]}), + ( + lambda value: _raised(lambda: ZarrV2ConsolidatedMetadata.from_json(value)), + {"zarr_consolidated_format": [1], "metadata": {"a/.zarray": [math.inf]}}, + ), + ( + lambda value: _raised(lambda: ZarrV3ArrayMetadata.from_key_value(value)), # pyright: ignore[reportArgumentType] + {"zarr.json": b"{"}, + ), +] + + +@pytest.mark.parametrize( + ("read", "value"), + READERS, + ids=[ + "check", + "read_configuration", + "resolve", + "fill-value-problems", + "validate-json", + "validate-metadata-field", + "validate-node-name", + "validate-node-path", + "validate-array", + "read-array", + "parse-array", + "array-from-json", + "validate-node", + "validate-group-and-what-it-holds", + "validate-array-v2", + "validate-group-v2", + "v2-consolidated-from-json", + "from-key-value", + ], +) +def test_every_problem_a_reader_finds_holds_what_the_value_handed_in_holds_there( + read: Callable[[object], Sequence[ValidationProblem]], value: object +) -> None: + # The value itself, not a copy a reader made on the way: the problems + # of a field, a pipeline, a chunk grid over the shape, and a document + # a group holds alike. A missing key holds none, and nor does what is + # not JSON, such as a store's bytes. + problems = read(value) + assert len(problems) != 0 + for problem in problems: + there = value_at(value, problem.loc) + assert problem.input is (there if is_canonical_json(there, finite=False) else UNSET), ( + problem + ) + if value is BAD_ARRAY: + assert {problem.loc for problem in problems} >= BAD_ARRAY_LOCS + + +class WideBits(TypedDict, closed=True): + bits: Annotated[int, Le(16)] + + +def test_a_problem_with_what_a_name_carries_is_the_field_s() -> None: + # `r24` carries `{"bits": 24}`, which this reader's own raw bits refuse: + # the problem is found at the field, where the name is what is there, + # and what was expected of `bits` is not expected of the name. + scope = CORE_AND_EXTENSIONS.extended_with(DataTypeDefinition(name="r*", configuration=WideBits)) + _, problems = resolve("r24", DataTypeDefinition, scope, ("data_type",)) + assert [(p.loc, p.input, dict(p.ctx)) for p in problems] == [(("data_type",), "r24", {})] + assert problems[0].message == "expected an integer <= 16, got 24" + + +class Ordered(TypedDict, closed=True): + order: Literal["F", "C"] + + +@pytest.mark.parametrize( + ("problems", "loc", "expected"), + [ + (check({"order": "Q"}, Ordered)[1], ("order",), ("C", "F")), + ( + validate_array_metadata_v3({**BAD_ARRAY, "zarr_format": 2}), + ("zarr_format",), + (3,), + ), + (validate_node_metadata_v3({"node_type": "dataset"}), ("node_type",), ("array", "group")), + (validate_array_metadata_v2({"order": "Q"}), ("order",), ("C", "F")), + ( + validate_group_metadata_v3( + { + "zarr_format": 3, + "node_type": "group", + "consolidated_metadata": { + "kind": "inline", + "must_understand": True, + "metadata": {}, + }, + } + ), + ("consolidated_metadata", "must_understand"), + (False,), + ), + ], + ids=["a-literal", "zarr-format", "node-type", "v2-order", "consolidated-must-understand"], +) +def test_a_value_outside_a_closed_set_is_told_the_set( + problems: Sequence[ValidationProblem], loc: tuple[str | int, ...], expected: tuple[object, ...] +) -> None: + # In the order the message lists them, as JSON writes them. + (problem,) = [problem for problem in problems if problem.loc == loc] + assert dict(problem.ctx) == {"expected": expected} + + +def test_error_a_ctx_of_another_mapping_type_is_checked() -> None: + # Only a problem's own `ctx`, checked when it was made, is taken as + # checked: any other mapping is, whatever its type. + with pytest.raises(TypeError, match="ctx is an object of JSON values"): + ValidationProblem( + (), "m", "invalid_value", ctx=cast("Any", MappingProxyType({"x": object()})) + ) + first = ValidationProblem((), "m", "invalid_value", ctx={"ge": 0}) + again = dataclasses.replace(first, loc=("a",)) + assert again.ctx == {"ge": 0} + assert again.ctx is first.ctx + + +def test_a_problem_at_the_first_element_holds_it() -> None: + # Checked against the literal, not `value_at`, which the reader uses. + (problem,) = with_input([ValidationProblem(("a", 0), "bad", "invalid_value")], {"a": [7]}) + assert problem.input == 7 + assert value_at([5, 6], (0,)) == 5 + read = resolve("bytes", DataTypeDefinition, CORE_AND_EXTENSIONS)[0] + (found,) = fill_value_problems(read, [-1]) + assert (found.loc, found.input) == ((0,), -1) + + +def test_a_problem_s_ctx_is_its_own_at_every_level() -> None: + # Copied when the problem is made, at every level, so nothing the + # caller does to what it handed in reaches the problem. + inner: dict[str, int] = {"a": 1} + problem = ValidationProblem((), "m", "invalid_value", ctx={"nested": inner}) + inner["a"] = 2 + assert problem.ctx["nested"] == {"a": 1} + + +def test_error_a_problem_not_below_where_a_reader_counts_from_is_refused() -> None: + # `within` relocates a problem from the document handed in to the one + # read; one located from another root would land in the wrong place. + problem = ValidationProblem(("a", "b"), "m", "invalid_value") + assert within((problem,), ("a",))[0].loc == ("b",) + with pytest.raises(TypeError, match="does not sit below"): + within((problem,), ("c",)) diff --git a/packages/zarr-metadata/tests/test_public_api.py b/packages/zarr-metadata/tests/test_public_api.py index cb88193103..7d95ef2cb4 100644 --- a/packages/zarr-metadata/tests/test_public_api.py +++ b/packages/zarr-metadata/tests/test_public_api.py @@ -43,17 +43,17 @@ def _group_rank(s: str) -> int: "JSONValue", # Category A' — metadata models (in-memory dataclasses over the documents) "ZarrV2ArrayMetadata", - "ZarrV2ArrayMetadataPartial", + "ZarrV2ArrayMetadataUpdate", "ZarrV3ArrayMetadata", - "ZarrV3ArrayMetadataPartial", + "ZarrV3ArrayMetadataUpdate", "ZarrV2GroupMetadata", - "ZarrV2GroupMetadataPartial", + "ZarrV2GroupMetadataUpdate", "ZarrV3GroupMetadata", - "ZarrV3GroupMetadataPartial", + "ZarrV3GroupMetadataUpdate", + "ZarrV3ConsolidatedMetadataInput", + "ZarrV3NodeMetadataInput", "ZarrV2ConsolidatedMetadata", "ZarrV3ConsolidatedMetadata", - "ZarrV3NamedConfig", - "ZarrV3MetadataField", "ValidationProblem", "MetadataValidationError", "ProblemKind", @@ -240,7 +240,7 @@ def test_all_is_grouped_and_unique() -> None: # Core document/model names: the format version comes first (`ZarrV2` / # `ZarrV3`), then the CamelCase entity, then an optional role suffix -# (`JSON`, `JSONPartial`, `Partial`, `StoreKey`) — validated loosely here +# (`JSON`, `JSONPartial`, `Partial`, `Reading`, `StoreKey`) — validated loosely here # because `JSON` decomposes into single-letter words under any strict # word-splitting regex. _CORE_NAME = re.compile(r"^ZarrV[23](?:[A-Z][a-z0-9]*)+$") @@ -286,6 +286,13 @@ def test_all_is_grouped_and_unique() -> None: "BloscCName", "BloscShuffle", "Chunk", + # The algebra of scopes: what a reading claims, and where two + # scopes disagree. + "ClaimKey", + "Claims", + "Conflict", + "Disagreements", + "ScopeConflictError", "CodecKind", "CodecSize", "Context", @@ -293,17 +300,23 @@ def test_all_is_grouped_and_unique() -> None: "Lengths", "Loc", "Nested", - "Resolution", - "Resolved", + "NodeName", + "NodePath", + # What a scope made of a field: `AcceptedField` by the definition that claims + # its name, `UnclaimedField`, or `RefusedField`; `ResolvedField` is the three. + "AcceptedField", + "RefusedField", + "ResolvedField", "Stage", "StorageClass", - "Unread", + "UnclaimedField", "CastOutOfRangeMode", "CastRoundingMode", "Endianness", "HexFloat16", "HexFloat32", "HexFloat64", + "JSONSchema", "JSONValue", "MetadataValidationError", "NumpyDatetime64", @@ -311,12 +324,15 @@ def test_all_is_grouped_and_unique() -> None: "NumpyTimedelta64", "ProblemKind", "RectilinearDimSpec", + # What a repair of a writer's bug changed, as `ValidationProblem` and + # `ProblemKind` are what a read found. + "Repair", + "RepairKind", "ScalarMap", "ScalarMapEntry", "ShardingIndexLocation", "Struct", "StructField", - "TypedDictKeys", "ValidationProblem", } ) diff --git a/packages/zarr-metadata/tests/test_typed_json.py b/packages/zarr-metadata/tests/test_typed_json.py index f225ba9eae..311d497395 100644 --- a/packages/zarr-metadata/tests/test_typed_json.py +++ b/packages/zarr-metadata/tests/test_typed_json.py @@ -7,6 +7,7 @@ from __future__ import annotations +import dataclasses import importlib import itertools import math @@ -15,10 +16,34 @@ import textwrap import types from collections.abc import Mapping -from typing import TYPE_CHECKING, Generic, Literal, NewType, NotRequired, TypeVar, cast +from typing import ( + TYPE_CHECKING, + Annotated, + Generic, + Literal, + NewType, + NotRequired, + Required, + TypeVar, + cast, +) +import pydantic import pytest -from typing_extensions import ReadOnly, TypeAliasType, TypedDict, is_typeddict +from annotated_types import ( + BaseMetadata, + Ge, + Gt, + Interval, + Le, + Len, + Lt, + MinLen, + MultipleOf, + Predicate, + Timezone, +) +from typing_extensions import Doc, ReadOnly, TypeAliasType, TypedDict, is_typeddict import zarr_metadata.v2 import zarr_metadata.v3 @@ -34,10 +59,16 @@ shape_of, typeddict_keys, ) -from zarr_metadata.model import ZarrV2ArrayMetadata, ZarrV3ArrayMetadata -from zarr_metadata.typed_json import check +from zarr_metadata.model import ( + ZarrV2ArrayMetadata, + ZarrV3ArrayMetadata, +) +from zarr_metadata.typed_json import ( + check, +) from zarr_metadata.v2.array import ZarrV2ArrayMetadataJSON, ZarrV2DataTypeMetadata from zarr_metadata.v3.array import ZarrV3ArrayMetadataJSON +from zarr_metadata.v3.data_type.float32 import Float32FillValue if TYPE_CHECKING: from _pytest.mark import ParameterSet @@ -171,7 +202,8 @@ def test_a_value_of_the_shape_reads_as_itself( (str, 1, (), "invalid_type"), (None, 0, (), "invalid_type"), (Literal["a"], "b", (), "invalid_value"), - (Literal[1], True, (), "invalid_value"), + (Literal[1], True, (), "invalid_type"), + (Literal["0.5"], 0.5, (), "invalid_type"), (tuple[int, ...], (1, "x"), (1,), "invalid_type"), (tuple[int, ...], 5, (), "invalid_type"), (int | str, 2.5, (), "invalid_type"), @@ -184,6 +216,7 @@ def test_a_value_of_the_shape_reads_as_itself( "invalid_type", ), (Blosc | Gzip, {"name": "zstd", "configuration": {}}, ("name",), "invalid_value"), + (Blosc | Gzip, {"name": 5, "configuration": {}}, ("name",), "invalid_type"), (Blosc | Gzip, {"configuration": {}}, ("name",), "missing_key"), (Triple | Pair, {"a": 1, "b": "x"}, ("b",), "invalid_type"), ], @@ -195,6 +228,7 @@ def test_a_value_of_the_shape_reads_as_itself( "zero-is-not-null", "literal-other", "literal-bool-is-not-int", + "literal-of-another-type", "element", "not-a-sequence", "no-branch", @@ -202,6 +236,7 @@ def test_a_value_of_the_shape_reads_as_itself( "not-json", "a-tag-picks-the-branch-that-reports", "a-tag-no-branch-has", + "a-tag-of-another-type", "a-tag-missing", "untagged-the-closest-branch-reports", ], @@ -871,13 +906,213 @@ def test_a_shape_is_what_a_union_dispatches_on(annotation: object, shape: str | assert shape_of(annotation) == shape +@pytest.mark.parametrize( + ("annotation", "value", "message"), + [ + (Blosc | Gzip, 3, "expected an object, got 3"), + (Literal["0.5"], 0.5, 'expected "0.5", got 0.5'), + (Literal["C", "F"], "Q", 'expected one of ["C", "F"], got "Q"'), + (tuple[str, ...], "nuclei", 'expected an array, got "nuclei"'), + (str, None, "expected a string, got null"), + ], + ids=["a-union-of-objects", "one-value", "values", "array", "null"], +) +def test_a_message_names_what_was_expected_and_shows_the_json_it_got( + annotation: object, value: object, message: str +) -> None: + assert [problem.message for problem in _read(annotation, value)[1]] == [message] + + def test_a_shape_is_described_as_a_message_would_name_it() -> None: - assert describe(Literal[0, "auto"]) == "one of ('auto', 0)" + assert describe(Literal[0, "auto"]) == 'one of ["auto", 0]' + assert describe(Literal["0.5"]) == '"0.5"' assert describe(int | None) == "an integer or null" + assert describe(Blosc | Gzip) == "an object" + assert describe(float | int | None) == "a number or null" + # An array's elements are named as many. + assert describe(tuple[int, ...]) == "an array of integers" + assert describe(tuple[int | None, ...]) == "an array of integers or nulls" + assert describe(tuple[tuple[str, ...], ...]) == "an array of arrays of strings" + assert describe(tuple[Literal["C", "F"], ...]) == 'an array of values in ["C", "F"]' + assert describe(int | tuple[int, ...]) == "an integer or an array of integers" + # A broader shape takes in a narrower one: the hex string a float32 + # is spelled as, the strings of its non-finite values. + assert describe(Float32FillValue) == "a number or a string" + assert describe(Literal["C", "F"] | None) == 'one of ["C", "F"] or null' assert describe(Width) == "an integer" assert describe(ZarrV2DataTypeMetadata) == "a ZarrV2DataTypeMetadata" +# --- constraints ------------------------------------------------------------ + +Digit = Annotated[int, Interval(ge=0, le=9)] + + +class Bounded(TypedDict, extra_items=Annotated[int, Ge(0)]): + level: Digit + shape: ReadOnly[NotRequired[tuple[Annotated[int, Ge(1)], ...]]] + name: Annotated[NotRequired[Annotated[str, "the inner note"]], Doc("the outer note")] + + +class Noted(TypedDict, extra_items=Annotated[object, "any JSON"]): + level: int + + +@pytest.mark.parametrize( + ("annotation", "value", "found"), + [ + (Digit, 9, []), + (Digit, 10, [((), "expected an integer in [0, 9], got 10", {"ge": 0, "le": 9})]), + (Digit, -1, [((), "expected an integer in [0, 9], got -1", {"ge": 0, "le": 9})]), + (Annotated[int, Gt(0)], 0, [((), "expected an integer > 0, got 0", {"gt": 0})]), + (Annotated[float, Lt(1)], 1, [((), "expected a number < 1, got 1", {"lt": 1})]), + (Annotated[float, Le(0.5)], 0.5, []), + ( + Annotated[float, Interval(gt=0, le=0.5)], + 0.0, + [((), "expected a number in (0, 0.5], got 0.0", {"gt": 0, "le": 0.5})], + ), + ( + Annotated[float, Interval(ge=0.0, le=9.0)], + 10, + [((), "expected a number in [0, 9], got 10", {"ge": 0, "le": 9})], + ), + (Annotated[int, Ge(True)], 0, [((), "expected an integer >= 1, got 0", {"ge": 1})]), + ( + tuple[Digit, ...], + [0, 10, 9], + [((1,), "expected an integer in [0, 9], got 10", {"ge": 0, "le": 9})], + ), + (Annotated[Width, Le(9)], 10, [((), "expected an integer <= 9, got 10", {"le": 9})]), + (Digit | str, "ten", []), + (Digit | str, 10, [((), "expected an integer in [0, 9], got 10", {"ge": 0, "le": 9})]), + (Digit, "nine", [((), 'expected an integer, got "nine"', {})]), + (Annotated[int, "a note", Doc("another")], -1, []), + ], + ids=[ + "within", + "above", + "below", + "exclusive-lower", + "exclusive-upper", + "inclusive-upper", + "half-open", + "an-integral-bound-is-an-integer", + "a-bool-bound-is-the-integer-it-equals", + "each-element-at-its-index", + "on-an-alias", + "a-branch-it-does-not-bound", + "the-branch-it-bounds", + "a-value-of-another-type-is-that-problem-alone", + "notes", + ], +) +def test_a_value_is_held_to_the_bounds_its_type_carries( + annotation: object, value: object, found: list[tuple[Loc, str, dict[str, JSONValue]]] +) -> None: + # annotated-types' bounds, as pydantic reads them: one problem for a + # value out of any, saying what the type admits, its bounds in the + # problem's `ctx`; and asked only of a value of the type. A bound is + # the number it equals, whichever of the equal ones Python's + # `Annotated` cache hands back. + _, problems = _read(annotation, value) + assert [(problem.loc, problem.message, dict(problem.ctx)) for problem in problems] == found + assert all(problem.kind == "invalid_value" for problem in problems if problem.ctx) + + +def test_a_key_s_type_keeps_the_metadata_its_annotation_carries() -> None: + # Qualifiers peeled, wherever they sit among the `Annotated` layers, + # and the metadata kept, an inner layer's first, as `Annotated` + # flattens nested layers; a note on `object` leaves the type open. + keys = typeddict_keys(Bounded) + assert keys.members["level"] == (Digit, True) + assert keys.members["shape"] == (tuple[Annotated[int, Ge(1)], ...], False) + assert keys.members["name"] == ( + Annotated[str, "the inner note", Doc("the outer note")], + False, + ) + assert keys.extra_items == Annotated[int, Ge(0)] + typed, problems = check({"level": 3, "shape": [0], "other": -1}, Bounded) + assert typed is None + assert [(problem.loc, problem.input, dict(problem.ctx)) for problem in problems] == [ + (("other",), -1, {"ge": 0}), + (("shape", 0), 0, {"ge": 1}), + ] + assert typeddict_keys(Noted).open + assert check({"level": 1, "more": [None]}, Noted) == ({"level": 1, "more": (None,)}, ()) + + +def test_error_a_bound_on_what_is_not_a_number() -> None: + with pytest.raises(TypeError, match="ge: a bound is on a number, and a string is not one"): + parser(Annotated[str, Ge(0)], no_leaf) + + +def test_error_a_bound_on_the_values_an_open_typeddict_takes() -> None: + class Loose(TypedDict, extra_items=Annotated[object, Ge(0)]): + level: int + + with pytest.raises(TypeError, match="ge: a bound is on a number, and a value is not one"): + parser(Loose, no_leaf) + + +@dataclasses.dataclass(frozen=True) +class Later(BaseMetadata): + """A constraint a later annotated-types might add.""" + + +@pytest.mark.parametrize( + "constraint", + [ + MinLen(1), + Len(1, 3), + MultipleOf(2), + Predicate(str.isdigit), + Timezone(None), + Later(), + pydantic.Field(ge=0), + pydantic.StringConstraints(pattern="^[a-z]+$"), + ], + ids=[ + "min-len", + "len", + "multiple-of", + "predicate", + "timezone", + "one-a-later-release-adds", + "pydantic-field", + "pydantic-string-constraints", + ], +) +def test_error_a_constraint_the_checker_does_not_read(constraint: object) -> None: + # A type the checker did not hold its values to would say what is not + # so, as a type pydantic reads and the checker does not would disagree. + with pytest.raises(TypeError, match="is not a constraint the checker reads"): + parser(Annotated[str, constraint], no_leaf) + + +@pytest.mark.parametrize( + "annotation", + [ + Annotated[Digit, Ge(1)], + Annotated[int, Ge(0), Gt(0)], + Annotated[int, Interval(le=9), Lt(5)], + ], + ids=["narrowing-an-alias", "gt-and-ge", "interval-and-lt"], +) +def test_error_a_second_bound_from_one_side(annotation: object) -> None: + # Pydantic reads the last one said; say one. + with pytest.raises(TypeError, match="is a second bound from (below|above)"): + parser(annotation, no_leaf) + + +@pytest.mark.parametrize( + "bound", ["0", math.nan, math.inf, None], ids=["string", "nan", "infinity", "null"] +) +def test_error_a_bound_that_is_not_a_finite_number(bound: object) -> None: + with pytest.raises(TypeError, match="a bound is a finite number"): + parser(Annotated[int, Ge(bound)], no_leaf) # pyright: ignore[reportArgumentType] + + # --- check: the public door ------------------------------------------------ _V3_DOCUMENT = dict(ZarrV3ArrayMetadata.create_default(shape=(4,)).to_json()) @@ -995,3 +1230,19 @@ def test_every_typeddict_the_package_declares_compiles(typeddict: type) -> None: # A declaration no parser reads would be a document type `check` could # not be asked about. assert parser_for(typeddict, no_leaf) is not None + + +def test_a_union_of_a_string_and_another_shape_names_both() -> None: + # "a string" takes in only the Literals of strings, not every other branch. + assert describe(str | int) == "a string or an integer" + assert describe(str | None) == "a string or null" + (problem,) = parser(str | int, no_leaf)(None, ())[1] + assert problem.message == "expected a string or an integer, got null" + + +class _Three(TypedDict, closed=True): + x: Annotated[Required[Annotated[ReadOnly[Annotated[int, "inner"]], "middle"]], "outer"] + + +def test_metadata_of_three_annotated_layers_is_kept_in_order() -> None: + assert typeddict_keys(_Three).members["x"] == (Annotated[int, "inner", "middle", "outer"], True) diff --git a/packages/zarr-metadata/tests/test_typed_json_properties.py b/packages/zarr-metadata/tests/test_typed_json_properties.py index 00a137f560..b228546baf 100644 --- a/packages/zarr-metadata/tests/test_typed_json_properties.py +++ b/packages/zarr-metadata/tests/test_typed_json_properties.py @@ -13,27 +13,31 @@ Values are JSON, their depth capped: drawn from the spec, so they conform; drawn at large, so most do not; or drawn from the spec and changed in one place, so they nearly do. On every one the checker has to -agree with the reference, and the two builds with each other. +agree with the reference, and the two builds with each other; so does +the JSON Schema written of the type, but for JSON Schema's own reading of +a number with no fraction, `1.0`, as the integer it equals. """ from __future__ import annotations import copy import itertools +import json import sys import types import typing from collections.abc import Mapping from dataclasses import dataclass, field -from typing import Annotated, Literal, NewType, NotRequired, Required, TypeAlias, Union, cast +from typing import Annotated, Any, Literal, NewType, NotRequired, Required, TypeAlias, Union, cast import pytest from hypothesis import HealthCheck, event, find, given, note, settings from hypothesis import strategies as st +from jsonschema import Draft202012Validator from typing_extensions import ReadOnly, TypeAliasType, TypedDict from zarr_metadata._common import JSONValue -from zarr_metadata._typed_json import Loc, Parsed, no_leaf, parser, typeddict_keys +from zarr_metadata._typed_json import Loc, Parsed, Schemas, no_leaf, parser, typeddict_keys # --- specs ----------------------------------------------------------------- @@ -664,6 +668,38 @@ def test_the_checker_agrees_with_the_reference(data: st.DataObject) -> None: assert conforms(typed, spec) +def _as_json_schema_reads(value: object) -> object: + """`value` as JSON Schema reads it: a number with no fraction is an integer, `1.0` the `1` it equals.""" + if isinstance(value, float) and value.is_integer(): + return int(value) + if isinstance(value, list): + return [_as_json_schema_reads(entry) for entry in cast("list[object]", value)] + if isinstance(value, dict): + entries = cast("dict[str, object]", value) + return {key: _as_json_schema_reads(entry) for key, entry in entries.items()} + return value + + +@_EXAMPLES +@given(st.data()) +def test_the_json_schema_agrees_with_the_reference(data: st.DataObject) -> None: + # The schema of a type accepts exactly the values of it, as the checker + # does, but that JSON Schema takes `1.0` for the integer it equals, which + # the checker does not. + spec = data.draw(specs(_DEPTH), label="spec") + build = Build(data.draw(st.sampled_from(_MODES), label="mode")) + annotation = build.annotation(spec) + note(build.text_of()) + schemas = Schemas() + schema = schemas.document(schemas.of(annotation)) + note(json.dumps(schema, indent=1)) + Draft202012Validator.check_schema(schema) + value = data.draw(_values(spec), label="value") + expected = conforms(_as_json_schema_reads(value), spec) + event("conforms" if expected else "does not conform") + assert Draft202012Validator(schema).is_valid(cast("Any", value)) == expected + + @_EXAMPLES @given(st.data()) def test_a_postponed_typeddict_reads_as_an_evaluated_one(data: st.DataObject) -> None: diff --git a/packages/zarr-metadata/tests/v2/test_codecs.py b/packages/zarr-metadata/tests/v2/test_codecs.py new file mode 100644 index 0000000000..28447f887c --- /dev/null +++ b/packages/zarr-metadata/tests/v2/test_codecs.py @@ -0,0 +1,127 @@ +"""Every v2 codec, read as its numcodecs configuration.""" + +from __future__ import annotations + +import json +from typing import Any + +import pytest + +from zarr_metadata.v2._definition import ZarrV2CodecDefinition +from zarr_metadata.v2.codec import V2_CODECS, ZarrV2CodecMetadata +from zarr_metadata.v3._definition import ( + canonical_of, +) +from zarr_metadata.v3.definition import ( + AcceptedField, + Context, + RefusedField, + UnclaimedField, + resolve, +) + +SCOPE = Context.of(*V2_CODECS) +Loc = tuple[str | int, ...] + +# numcodecs 0.16.5 `get_config()` of a default instance, where one exists. +EXAMPLES: dict[str, tuple[dict[str, Any], ...]] = { + # -1 is zlib's default-level constant, which numcodecs writes as given. + "zlib": ({"id": "zlib", "level": 1}, {"id": "zlib"}, {"id": "zlib", "level": -1}), + "gzip": ({"id": "gzip", "level": 9}, {"id": "gzip", "level": -1}), + "bz2": ({"id": "bz2", "level": 1},), + "lzma": ({"id": "lzma", "format": 1, "check": -1, "preset": None, "filters": None},), + "blosc": ( + {"id": "blosc", "cname": "lz4", "clevel": 5, "shuffle": 1, "blocksize": 0}, + {"id": "blosc", "clevel": 5}, + {"id": "blosc", "cname": "zstd", "clevel": 3, "shuffle": 2, "blocksize": 0, "typesize": 4}, + ), + "zstd": ({"id": "zstd", "level": 0, "checksum": False},), + "lz4": ({"id": "lz4", "acceleration": 1},), + "shuffle": ({"id": "shuffle", "elementsize": 4},), + "delta": ({"id": "delta", "dtype": " None: + """Every definition in `V2_CODECS` is exercised below.""" + assert {definition.name for definition in V2_CODECS} == set(EXAMPLES) + + +@pytest.mark.parametrize( + ("name", "field"), CASES, ids=[f"{n}:{i}" for i, (n, _) in enumerate(CASES)] +) +def test_every_example_reads_and_is_written_back_as_it_was( + name: str, field: ZarrV2CodecMetadata +) -> None: + """A numcodecs configuration reads by the definition its id names, with no problem, and is written back as it was: the id beside the parameters, which have one spelling.""" + resolved, problems = resolve(field, ZarrV2CodecDefinition, SCOPE) + assert isinstance(resolved, AcceptedField) + assert resolved.definition.name == name + assert problems == () + assert resolved.to_json() == field + assert canonical_of(resolved, problems) == field + + +@pytest.mark.parametrize( + "field", + [{"id": "categorize", "labels": ["a"]}, {"id": "pickle"}, {"id": "n5_wrapper", "inner": 1}], +) +def test_an_id_the_package_does_not_model_is_unclaimed(field: dict[str, Any]) -> None: + """A codec id nothing in scope claims reads as `UnclaimedField`, its parameters kept and unjudged, as a v3 extension nothing claims is.""" + resolved, problems = resolve(field, ZarrV2CodecDefinition, SCOPE) + assert isinstance(resolved, UnclaimedField) + assert problems == () + assert json.dumps(resolved.to_json()) == json.dumps(field) + + +@pytest.mark.parametrize( + ("field", "at", "kind"), + [ + ({"id": "zlib", "level": 10}, ("c", "level"), "invalid_value"), + ({"id": "gzip", "level": -2}, ("c", "level"), "invalid_value"), + ({"id": ""}, ("c", "id"), "invalid_value"), + ({"id": "bz2", "level": 0}, ("c", "level"), "invalid_value"), + ({"id": "blosc", "cname": "brotli"}, ("c", "cname"), "invalid_value"), + ({"id": "blosc", "clevel": -1}, ("c", "clevel"), "invalid_value"), + ({"id": "blosc", "shuffle": 3}, ("c", "shuffle"), "invalid_value"), + ({"id": "blosc", "blocksize": -1}, ("c", "blocksize"), "invalid_value"), + ({"id": "zstd", "level": 23}, ("c", "level"), "invalid_value"), + ({"id": "shuffle", "elementsize": 0}, ("c", "elementsize"), "invalid_value"), + ({"id": "delta"}, ("c", "dtype"), "missing_key"), + ({"id": "delta", "dtype": "float32"}, ("c", "dtype"), "invalid_value"), + ( + {"id": "astype", "encode_dtype": " None: + """A parameter out of its range, of the wrong type, missing when numcodecs has no default, a dtype parameter that is no typestr, or a key no codec declares is reported at the parameter, beside the field.""" + resolved, problems = resolve(field, ZarrV2CodecDefinition, SCOPE, ("c",)) + assert [(p.loc, p.kind) for p in problems] == [(at, kind)] + assert isinstance(resolved, AcceptedField if kind == "unknown_key" else RefusedField) diff --git a/packages/zarr-metadata/tests/v2/test_data_types.py b/packages/zarr-metadata/tests/v2/test_data_types.py new file mode 100644 index 0000000000..880b943505 --- /dev/null +++ b/packages/zarr-metadata/tests/v2/test_data_types.py @@ -0,0 +1,305 @@ +"""Every v2 data type, read as its family: which typestrs it takes, and which fill values.""" + +from __future__ import annotations + +from typing import TYPE_CHECKING + +import pytest + +from zarr_metadata.v2._definition import ZarrV2DataTypeDefinition +from zarr_metadata.v2.data_type import V2_DATA_TYPES +from zarr_metadata.v3._definition import ( + canonical_of, + fields_of, +) +from zarr_metadata.v3.definition import ( + AcceptedField, + Context, + RefusedField, + UnclaimedField, + canonical_fill_value, + fill_value_problems, + resolve, +) + +if TYPE_CHECKING: + from zarr_metadata import JSONValue + +SCOPE = Context.of(*V2_DATA_TYPES) + +Loc = tuple[str | int, ...] + + +def _read(value: object) -> tuple[type, list[tuple[Loc, str]]]: + resolved, problems = resolve(value, ZarrV2DataTypeDefinition, SCOPE, ("dtype",)) + return type(resolved), [(p.loc, p.kind) for p in problems] + + +@pytest.mark.parametrize( + ("value", "family", "canonical"), + [ + ("|b1", "bool", "|b1"), + ("i4", "int", ">i4"), + ("u8", "uint", ">u8"), + ("f8", "float", ">f8"), + ("c16", "complex", ">c16"), + ("|S0", "bytes", "|S0"), + ("|S12", "bytes", "|S12"), + ("U0", "str", ">U0"), + ("|V8", "void", "|V8"), + ("|O", "object", "|O"), + ("M8[10s]", "datetime64", ">M8[10s]"), + (" None: + """Each typestr of a family reads by the family's definition, and its simplest spelling is the typestr NumPy writes: `|` for a type of one byte or no byte order, `us` for a microsecond unit; a records array reads as `struct`, each record's type read too.""" + resolved, problems = resolve(value, ZarrV2DataTypeDefinition, SCOPE) + assert isinstance(resolved, AcceptedField) + assert resolved.definition.name == family + assert problems == () + assert canonical_of(resolved, problems) == canonical + + +@pytest.mark.parametrize("value", [" None: + """A typestr whose type code the v2 spec does not list is filed under itself, which nothing in scope claims: read as `UnclaimedField`, with no problem, and written back as it was.""" + resolved, problems = resolve(value, ZarrV2DataTypeDefinition, SCOPE) + assert isinstance(resolved, UnclaimedField) + assert problems == () + assert resolved.to_json() == value + + +@pytest.mark.parametrize( + ("value", "at"), + [ + ("float32", ("dtype",)), + (" None: + """A typestr of a size, byte order or unit its family does not take is refused, the problem at the field; a struct's problems sit under `fields`, at the record and its position, and a record's type that is refused is reported there while the struct still reads.""" + kind, problems = _read(value) + assert kind is RefusedField or "fields" in at + assert next(loc for loc, _ in problems) == at + + +@pytest.mark.parametrize( + ("dtype", "value", "canonical"), + [ + ("|b1", True, True), + ("|b1", None, None), + (" None: + """Each family takes the fill values the v2 spec gives it, and `null` always: booleans, integers in the type's range, numbers or the named non-finite strings for floats and each complex component, base64 for bytes, void and struct, a string for str, any JSON for object, ticks or `NaT` for times. The canonical spelling makes an integer written for a float a float, and `NaT` of its ticks.""" + resolved, _ = resolve(dtype, ZarrV2DataTypeDefinition, SCOPE) + assert fill_value_problems(resolved, value) == () + assert canonical_fill_value(resolved, value) == canonical + + +@pytest.mark.parametrize( + ("dtype", "value", "at"), + [ + ("|b1", 1, ()), + (" None: + """A fill value outside the shape or range its family takes is reported, at the value or at the component that is wrong: a v2 float takes no hex string, and a struct's fill value is base64 of the record, not an object of fields.""" + resolved, _ = resolve(dtype, ZarrV2DataTypeDefinition, SCOPE) + problems = fill_value_problems(resolved, value, ("fill_value",)) + assert len(problems) != 0 + assert problems[0].loc == ("fill_value", *at) + + +def test_every_family_has_an_example() -> None: + """Every definition in `V2_DATA_TYPES` is exercised above, so none is filed untested.""" + names = {definition.name for definition in V2_DATA_TYPES} + assert names == { + "bool", + "int", + "uint", + "float", + "complex", + "bytes", + "str", + "void", + "datetime64", + "timedelta64", + "object", + "struct", + } + + +def test_a_struct_record_type_sits_where_the_document_writes_it() -> None: + """`fields_of` places a record's type under the struct's own configuration location, `("dtype", "fields", 0, 1)`, not under a `configuration` key a v2 document does not have.""" + resolved, _ = resolve([["a", " None: + """A number written for a float fill value that no float of the dtype's width holds -- ` None: + """A void or struct fill value decodes to exactly the item's size -- a record's the sum of its fields', each times the product of its subarray shape, nested structs included -- and a byte string's to at most it, as NumPy pads a shorter one; a record holding a field of no fixed size, `|O`, is left unjudged. zarr-python 3 fails to open an array whose fill value is of another size.""" + from zarr_metadata.model import ZarrV2ArrayMetadata, validate_array_metadata_v2 + + document = { + **ZarrV2ArrayMetadata.create_default(shape=(4,), chunks=(2,)).to_json(), + "dtype": dtype, + "fill_value": fill_value, + } + found = validate_array_metadata_v2(document) + assert (len(found) == 0) is accepted, [p.message for p in found] + if not accepted: + assert [(p.loc, p.kind) for p in found] == [(("fill_value",), "invalid_value")] diff --git a/packages/zarr-metadata/tests/v2/test_definition.py b/packages/zarr-metadata/tests/v2/test_definition.py new file mode 100644 index 0000000000..7e8332c3eb --- /dev/null +++ b/packages/zarr-metadata/tests/v2/test_definition.py @@ -0,0 +1,52 @@ +"""The public door to v2 definitions: the scope zarr-python 2.x reads in, and one field read in it.""" + +from __future__ import annotations + +import pickle + +import pytest + +from zarr_metadata.v2.definition import ( + CORE_V2, + V2_CODECS, + V2_DATA_TYPES, + AcceptedField, + Context, + UnclaimedField, + ZarrV2CodecDefinition, + ZarrV2DataTypeDefinition, + resolve_codec_v2, + resolve_dtype_v2, +) +from zarr_metadata.v3.codec.gzip import GZIP_CODEC +from zarr_metadata.v3.definition import ( + CORE_AND_EXTENSIONS, + CodecDefinition, +) + + +def test_core_v2_files_every_v2_definition_apart_from_v3() -> None: + """`CORE_V2` files the 12 data types and 21 codecs by their v2 kinds, and is a scope of format 2: a v2 `gzip` and a v3 `gzip` are two definitions of two kinds, and no scope files both, since `Context.joined` refuses two formats.""" + assert set(CORE_V2.definitions()) == {*V2_DATA_TYPES, *V2_CODECS} + assert CORE_V2.format == 2 + assert CORE_V2.claimant(ZarrV2CodecDefinition, "gzip") is not GZIP_CODEC + assert CORE_V2.claimant(ZarrV2CodecDefinition, "gzip") is not None + assert CORE_V2.claimant(CodecDefinition, "gzip") is None + assert CORE_AND_EXTENSIONS.claimant(CodecDefinition, "gzip") is GZIP_CODEC + assert CORE_V2.claimant(ZarrV2DataTypeDefinition, " None: + """`resolve_dtype_v2` and `resolve_codec_v2` read one field in `CORE_V2` when no scope is given, and in the scope given otherwise, prefixing every problem with `loc`.""" + dtype, problems = resolve_dtype_v2("i8", "i", {"byteorder": ">", "itemsize": 8}), + ("|b1", "b", {"byteorder": "|", "itemsize": 1}), + ("|S12", "S", {"byteorder": "|", "itemsize": 12}), + (" None: + """A NumPy typestr splits into its type code and what the name carries: the byte order, the item size when written, and a time unit with its multiplier when bracketed.""" + assert parse_typestr(name) == (code, carried) + + +@pytest.mark.parametrize( + "name", + [ + "float32", + "f4", + " None: + """A string without a byte order, a type code and a size in that order is not a typestr.""" + assert parse_typestr(name) is None + + +@pytest.mark.parametrize( + ("name", "filed", "carried"), + [ + (" None: + """A typestr is filed under its family and carries its byte order, size and unit; `struct` is filed as itself; a typestr of a type code the spec does not list is filed under itself, so a scope leaves it unclaimed; a family name is filed by nothing, since no document writes it.""" + assert ZarrV2DataTypeDefinition.spelled(name) == (filed, carried) + + +@pytest.mark.parametrize( + ("value", "name", "configuration"), + [ + (" None: + """A dtype field is a string, which is its name, or an array of field records, which is a `struct` whose configuration is the records; anything else names nothing.""" + assert ZarrV2DataTypeDefinition.named_configuration(value) == (name, configuration, ()) + + +@pytest.mark.parametrize( + ("value", "problems"), + [ + (" None: + """The envelope of a dtype field holds when it is a typestr or an array of records; a string that is neither is reported as an invalid value, naming the typestr form, and anything else as an invalid type.""" + found = ZarrV2DataTypeDefinition.envelope_problems(value) + assert [(p.loc, p.kind) for p in found] == problems + if value == "float32": + assert "typestr" in found[0].message + + +def test_a_v2_dtype_writes_back_as_a_string_or_records() -> None: + """`envelope_json` writes a dtype as the string that carries it, and a struct as its records; its name and configuration both sit at the field.""" + assert ZarrV2DataTypeDefinition.envelope_json(" None: + """A codec field is an object whose `id` names it and whose other members are its parameters; one without a string `id`, or not an object, names nothing, and the envelope says why.""" + assert ZarrV2CodecDefinition.named_configuration(value) == (name, configuration, ()) + assert [(p.loc, p.kind) for p in ZarrV2CodecDefinition.envelope_problems(value)] == problems + + +def test_a_v2_codec_writes_back_with_its_id_beside_its_parameters() -> None: + """`envelope_json` writes a codec as `{"id": name, **parameters}`: its parameters at the field and its name under `id`.""" + assert ZarrV2CodecDefinition.envelope_json("zlib", {"level": 1}) == {"id": "zlib", "level": 1} + assert ZarrV2CodecDefinition.configuration_loc(("compressor",)) == ("compressor",) + assert ZarrV2CodecDefinition.name_loc(("compressor",)) == ("compressor", "id") + assert ZarrV2CodecDefinition.spelled("zlib") == ("zlib", None) + + +def test_error_a_v2_definition_is_named_as_its_format_names_one() -> None: + """A v2 codec definition with an empty id, and a v2 data type definition named by a typestr a document would write rather than a family, are refused when built.""" + with pytest.raises(TypeError, match="so no document names it"): + ZarrV2CodecDefinition(name="", configuration=EmptyConfiguration) + with pytest.raises(TypeError, match="is how a document writes"): + ZarrV2DataTypeDefinition(name=" bytes codecs", "invalid_value" ) @@ -194,6 +209,10 @@ def acme_paired_rules( ) +INT8: Final = AcceptedField(json="int8", name="int8", definition=INT8_DATA_TYPE, configuration={}) +"""An `int8` field as a scope reads it: by its definition, holding nothing inside.""" + + def _locs(problems: tuple[ValidationProblem, ...]) -> list[tuple[tuple[str | int, ...], str]]: return [(found.loc, found.kind) for found in problems] @@ -213,64 +232,279 @@ def test_a_read_field_keeps_the_fields_it_read_inside() -> None: }, } resolved, _ = resolve(shard, CodecDefinition, CORE_AND_EXTENSIONS) + assert isinstance(resolved, AcceptedField) assert {loc: inner.json for loc, inner in resolved.nested.items()} == { ("codecs", 0): "bytes", ("index_codecs", 0): "bytes", ("index_codecs", 1): "crc32c", } - cast = {"name": "cast_value", "configuration": {"data_type": "int8"}} - resolved, _ = resolve(cast, CodecDefinition, CORE_AND_EXTENSIONS) - assert resolved.nested[("data_type",)].definition is INT8_DATA_TYPE - # A field holding none, or one that was not read, has nothing inside. - assert resolve("int8", DataTypeDefinition, CORE)[0].nested == {} - unread = {"name": "cast_value", "configuration": {"data_type": "int8", "rounding": 1}} - assert resolve(unread, CodecDefinition, CORE_AND_EXTENSIONS)[0].nested == {} + cast_value = {"name": "cast_value", "configuration": {"data_type": "int8"}} + resolved, _ = resolve(cast_value, CodecDefinition, CORE_AND_EXTENSIONS) + assert isinstance(resolved, AcceptedField) + assert resolved.nested[("data_type",)] == INT8 + # A field holding none has nothing inside, and one nothing in scope + # claims holds no field at all; one its check or rules refuse keeps + # what it read. + assert resolve("int8", DataTypeDefinition, CORE)[0] == INT8 + unclaimed = {"name": "acme.cast", "configuration": {"data_type": "int8"}} + assert isinstance(resolve(unclaimed, CodecDefinition, CORE_AND_EXTENSIONS)[0], UnclaimedField) + refused = {"name": "cast_value", "configuration": {"data_type": "int8", "rounding": 1}} + resolved, _ = resolve(refused, CodecDefinition, CORE_AND_EXTENSIONS) + assert isinstance(resolved, RefusedField) + assert resolved.nested[("data_type",)] == INT8 @pytest.mark.parametrize( - ("field", "kind", "resolution", "configuration", "problems"), + ("field", "kind", "name"), + [ + ("int8", DataTypeDefinition, "int8"), + ({"name": "gzip", "configuration": {"level": 1}}, CodecDefinition, "gzip"), + # Raw bits: the name written, not the one its definition is filed under. + ("r16", DataTypeDefinition, "r16"), + ({"name": "acme.codec"}, CodecDefinition, "acme.codec"), + ({"configuration": {}}, CodecDefinition, None), + (5, CodecDefinition, None), + ], + ids=["bare-name", "object", "raw-bits", "out-of-scope", "no-name", "not-a-field"], +) +def test_a_reading_says_the_name_the_field_was_written_with( + field: JSONValue, kind: type[Definition[Any]], name: str | None +) -> None: + assert resolve(field, kind, CORE_AND_EXTENSIONS)[0].name == name + + +KINDLESS: Final = Definition(name="acme.kindless", configuration=Empty) +"""A definition of no kind, which no scope files.""" + + +@pytest.mark.parametrize( + ("built", "read_as"), + [ + (INT8, DataTypeDefinition), + # A name read by the definition filed under another, as raw bits are. + ( + AcceptedField( + json="r16", name="r16", definition=RAW_BYTES_DATA_TYPE, configuration={"bits": 16} + ), + DataTypeDefinition, + ), + # Type arguments dropped, as `resolve` drops them. + ( + UnclaimedField(json="acme.t", name="acme.t", read_as=DataTypeDefinition[Any]), + DataTypeDefinition, + ), + ( + RefusedField( + json="int8", + name="int8", + read_as=DataTypeDefinition[Any], + definition=INT8_DATA_TYPE, + ), + DataTypeDefinition, + ), + # Not a field, so named nothing and claimed by nothing. + (RefusedField(json=5, name=None, read_as=CodecDefinition), CodecDefinition), + ], + ids=["read", "raw-bits", "unclaimed", "refused", "refused-nameless"], +) +def test_a_field_built_by_hand_is_of_the_kind_it_says( + built: AcceptedField[Any] | UnclaimedField | RefusedField[Any], read_as: type[Definition[Any]] +) -> None: + assert built.read_as is read_as + + +@pytest.mark.parametrize("definition", [None, KINDLESS], ids=["none", "kindless"]) +def test_error_a_field_read_by_hand_by_what_is_not_a_definition_of_a_kind( + definition: Definition[Any] | None, +) -> None: + with pytest.raises(TypeError, match="a field read is read by a definition of a kind"): + AcceptedField( + json="acme.kindless", + name="acme.kindless", + definition=definition, # pyright: ignore[reportArgumentType] + configuration={}, + ) + + +@pytest.mark.parametrize("name", ["int16", "r16"]) +def test_error_a_field_read_by_hand_by_a_definition_filed_under_another_name(name: str) -> None: + with pytest.raises( + TypeError, match=f"a field named '{name}' is read by the definition filed under it" + ): + AcceptedField(json=name, name=name, definition=INT8_DATA_TYPE, configuration={}) + + +def test_error_a_field_refused_by_hand_by_a_definition_of_another_kind() -> None: + with pytest.raises( + TypeError, match="read as a DataTypeDefinition is read by one, got CodecDefinition" + ): + RefusedField(json="gzip", name="gzip", read_as=DataTypeDefinition, definition=GZIP_CODEC) + + +@pytest.mark.parametrize("name", ["int16", None]) +def test_error_a_field_refused_by_hand_by_a_definition_filed_under_another_name( + name: str | None, +) -> None: + with pytest.raises( + TypeError, match=f"a field named {name!r} is read by the definition filed under it" + ): + RefusedField(json=name, name=name, read_as=DataTypeDefinition, definition=INT8_DATA_TYPE) + + +@pytest.mark.parametrize( + "build", + [ + lambda: UnclaimedField(json="acme.t", name="acme.t", read_as=Definition), + lambda: RefusedField(json="acme.t", name="acme.t", read_as=Definition), + ], + ids=["unclaimed", "refused"], +) +def test_error_a_field_built_by_hand_of_no_kind(build: Callable[[], object]) -> None: + with pytest.raises(TypeError, match="is not a kind of metadata"): + build() + + +def test_error_a_field_nothing_claims_built_by_hand_without_a_name() -> None: + with pytest.raises(TypeError, match="a field nothing in scope claims is named"): + UnclaimedField(json=5, name=None, read_as=CodecDefinition) # pyright: ignore[reportArgumentType] + + +@pytest.mark.parametrize( + ("field", "kind", "written"), + [ + # A data type with nothing to configure by its bare name, however + # it was written; every other field an object. + ("int8", DataTypeDefinition, "int8"), + ( + {"name": "int8", "configuration": {}, "must_understand": True}, + DataTypeDefinition, + "int8", + ), + ("crc32c", CodecDefinition, {"name": "crc32c"}), + ("default", ChunkKeyEncodingDefinition, {"name": "default"}), + ( + {"name": "regular", "configuration": {"chunk_shape": [2, 3]}}, + ChunkGridDefinition, + {"name": "regular", "configuration": {"chunk_shape": (2, 3)}}, + ), + # Each field a configuration holds written the same way. + ( + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [2], + "codecs": ["bytes"], + "index_codecs": ["bytes", {"name": "crc32c", "configuration": {}}], + }, + }, + CodecDefinition, + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": (2,), + "codecs": ({"name": "bytes"},), + "index_codecs": ({"name": "bytes"}, {"name": "crc32c"}), + }, + }, + ), + ( + {"name": "cast_value", "configuration": {"data_type": {"name": "int8"}}}, + CodecDefinition, + {"name": "cast_value", "configuration": {"data_type": "int8"}}, + ), + # One nothing in scope claims inside it too; one refused as written. + ( + { + "name": "acme.stack", + "configuration": { + "codecs": [ + "zfpy", + {"name": "gzip", "configuration": {"level": 12}, "must_understand": True}, + ] + }, + }, + CodecDefinition, + { + "name": "acme.stack", + "configuration": { + "codecs": ( + {"name": "zfpy"}, + {"name": "gzip", "configuration": {"level": 12}, "must_understand": True}, + ) + }, + }, + ), + # A name that carries its configuration is written alone. + ("r16", DataTypeDefinition, "r16"), + ({"name": "r16", "configuration": {}}, DataTypeDefinition, "r16"), + # Nothing in scope claims it: its configuration as written. + ("zfpy", CodecDefinition, {"name": "zfpy"}), + ( + {"name": "zfpy", "configuration": {"x": [1]}}, + CodecDefinition, + {"name": "zfpy", "configuration": {"x": (1,)}}, + ), + ({"name": "acme.t", "configuration": {}}, DataTypeDefinition, "acme.t"), + ], + ids=[ + "data-type", + "data-type-spelled-out", + "codec", + "chunk-key-encoding", + "configured", + "nested", + "nested-data-type", + "nested-unclaimed-and-refused", + "raw-bits", + "raw-bits-spelled-out", + "unclaimed-codec", + "unclaimed-configured", + "unclaimed-data-type", + ], +) +def test_a_field_is_written_as_every_reader_takes_it( + field: JSONValue, kind: type[Definition[Any]], written: JSONValue +) -> None: + resolved, _ = resolve(field, kind, SCOPE) + assert isinstance(resolved, (AcceptedField, UnclaimedField)) + assert resolved.to_json() == written + + +@pytest.mark.parametrize( + ("field", "kind", "variant", "configuration", "problems"), [ ( {"name": "gzip", "configuration": {"level": 5}}, CodecDefinition, - "read", + AcceptedField, {"level": 5}, [], ), - ("crc32c", CodecDefinition, "read", {}, []), - ({"name": "crc32c"}, CodecDefinition, "read", {}, []), - ({"name": "crc32c", "configuration": {}}, CodecDefinition, "read", {}, []), - ({"name": "crc32c", "must_understand": True}, CodecDefinition, "read", {}, []), - ("bytes", CodecDefinition, "read", {}, []), + ("crc32c", CodecDefinition, AcceptedField, {}, []), + ({"name": "crc32c"}, CodecDefinition, AcceptedField, {}, []), + ({"name": "crc32c", "configuration": {}}, CodecDefinition, AcceptedField, {}, []), + ({"name": "crc32c", "must_understand": True}, CodecDefinition, AcceptedField, {}, []), + ("bytes", CodecDefinition, AcceptedField, {}, []), ( {"name": "bytes", "configuration": {"endian": "big"}}, CodecDefinition, - "read", + AcceptedField, {"endian": "big"}, [], ), ( {"name": "regular", "configuration": {"chunk_shape": [2, 3]}}, ChunkGridDefinition, - "read", + AcceptedField, {"chunk_shape": (2, 3)}, [], ), - # An extent of 0 is right on a dimension of length 0, which only the - # array's shape can tell. - ( - {"name": "regular", "configuration": {"chunk_shape": [0, 3]}}, - ChunkGridDefinition, - "read", - {"chunk_shape": (0, 3)}, - [], - ), # An unknown key is survivable: reported, left out of the # configuration, and the field still read. ( {"name": "gzip", "configuration": {"level": 5, "extra": 1}}, CodecDefinition, - "read", + AcceptedField, {"level": 5}, [(("configuration", "extra"), "unknown_key")], ), @@ -278,19 +512,19 @@ def test_a_read_field_keeps_the_fields_it_read_inside() -> None: ( {"name": "acme.bounded", "configuration": {"level": 5, "windw": 10}}, CodecDefinition, - "read", + AcceptedField, {"level": 5}, [(("configuration", "windw"), "unknown_key")], ), # A key that is not required, written as a string under postponed # annotations, and every key of a `total=False` TypedDict. - ("acme.level", CodecDefinition, "read", {}, []), - ("acme.total", CodecDefinition, "read", {}, []), + ("acme.level", CodecDefinition, AcceptedField, {}, []), + ("acme.total", CodecDefinition, AcceptedField, {}, []), # A kind with type arguments is that kind. ( {"name": "gzip", "configuration": {"level": 5}}, CodecDefinition[Any], - "read", + AcceptedField, {"level": 5}, [], ), @@ -300,7 +534,7 @@ def test_a_read_field_keeps_the_fields_it_read_inside() -> None: "configuration": {"label": "a", "children": [{"label": "b", "children": []}]}, }, CodecDefinition, - "read", + AcceptedField, {"label": "a", "children": ({"label": "b", "children": ()},)}, [], ), @@ -312,26 +546,28 @@ def test_a_read_field_keeps_the_fields_it_read_inside() -> None: "configuration": {"fallback": {"codec": {"name": "gzip"}, "note": "x"}}, }, CodecDefinition, - "read", + AcceptedField, {"fallback": {"codec": {"name": "gzip"}, "note": "x"}}, [], ), - # `extra_items` of a field alias: every other key holds a codec. + # `extra_items` of a field alias: every other key holds a codec, + # held as a document writes it. ( {"name": "acme.routes", "configuration": {"fast": "crc32c"}}, CodecDefinition, - "read", - {"fast": "crc32c"}, + AcceptedField, + {"fast": {"name": "crc32c"}}, [], ), # Nothing in scope claims it: left unjudged, not refused. - ({"name": "zfpy", "configuration": {"x": 1}}, CodecDefinition, "out_of_scope", None, []), - # A nested field is read in the same scope; one out of scope is left be. + ({"name": "zfpy", "configuration": {"x": 1}}, CodecDefinition, UnclaimedField, None, []), + # A nested field is read in the same scope, and held as a document + # writes it; one out of scope is left be. ( {"name": "acme.stack", "configuration": {"codecs": ["crc32c", "zfpy"]}}, CodecDefinition, - "read", - {"codecs": ("crc32c", "zfpy")}, + AcceptedField, + {"codecs": ({"name": "crc32c"}, {"name": "zfpy"})}, [], ), ], @@ -344,7 +580,6 @@ def test_a_read_field_keeps_the_fields_it_read_inside() -> None: "bytes-bare", "bytes-endian", "regular-grid", - "regular-grid-zero-extent", "unknown-key", "unknown-key-before-the-rules", "not-required-postponed", @@ -360,13 +595,15 @@ def test_a_read_field_keeps_the_fields_it_read_inside() -> None: def test_a_field_is_read_in_scope( field: object, kind: type[Definition[Any]], - resolution: str, + variant: type, configuration: dict[str, object] | None, problems: list[tuple[tuple[str | int, ...], str]], ) -> None: resolved, found = resolve(field, kind, SCOPE) - assert resolved.resolution == resolution - assert resolved.configuration == configuration + assert type(resolved) is variant + assert ( + resolved.configuration if isinstance(resolved, AcceptedField) else None + ) == configuration assert _locs(found) == problems @@ -374,8 +611,8 @@ def test_error_a_rule_refuses_a_value() -> None: resolved, found = resolve( {"name": "gzip", "configuration": {"level": 12}}, CodecDefinition, SCOPE ) - assert resolved.resolution == "invalid" - assert resolved.configuration is None + assert isinstance(resolved, RefusedField) + assert resolved.definition is GZIP_CODEC assert _locs(found) == [(("configuration", "level"), "invalid_value")] @@ -383,20 +620,25 @@ def test_error_a_member_of_the_wrong_type_is_not_asked_of_the_rules() -> None: resolved, found = resolve( {"name": "gzip", "configuration": {"level": "5"}}, CodecDefinition, SCOPE ) - assert resolved.resolution == "invalid" + assert isinstance(resolved, RefusedField) assert _locs(found) == [(("configuration", "level"), "invalid_type")] def test_error_a_required_configuration_is_missing() -> None: resolved, found = resolve("gzip", CodecDefinition, SCOPE) - assert resolved.resolution == "invalid" + assert isinstance(resolved, RefusedField) assert _locs(found) == [(("configuration",), "missing_key")] def test_error_the_configuration_is_not_an_object() -> None: - # Unread, and still claimed by the definition its name names. + # RefusedField, and still claimed by the definition its name names. resolved, found = resolve({"name": "gzip", "configuration": 5}, CodecDefinition, SCOPE) - assert (resolved.resolution, resolved.definition) == ("invalid", GZIP_CODEC) + assert resolved == RefusedField( + json={"name": "gzip", "configuration": 5}, + name="gzip", + read_as=CodecDefinition, + definition=GZIP_CODEC, + ) assert [found.loc for found in found] == [("configuration",)] @@ -404,7 +646,7 @@ def test_error_must_understand_false_is_refused() -> None: # The envelope's problem, reported with the field; the configuration # was read, so a later layer can still judge the codec. resolved, found = resolve({"name": "crc32c", "must_understand": False}, CodecDefinition, SCOPE) - assert resolved.resolution == "read" + assert isinstance(resolved, AcceptedField) assert _locs(found) == [(("must_understand",), "invalid_value")] @@ -492,22 +734,29 @@ def test_error_null_is_not_a_field() -> None: # not a verdict. Read as a field or checked as a configuration, it is # a value of the wrong type. resolved, found = resolve(None, CodecDefinition, SCOPE) - assert (resolved.resolution, _locs(found)) == ("invalid", [((), "invalid_type")]) - configuration, found = GZIP_CODEC.judge(None) + assert (resolved, _locs(found)) == ( + RefusedField(json=None, name=None, read_as=CodecDefinition), + [((), "invalid_type")], + ) + configuration, found = GZIP_CODEC.read_configuration(None) assert (configuration, _locs(found)) == (None, [((), "invalid_type")]) def test_error_a_value_that_is_not_json() -> None: + # Held as `UNSET`, not JSON, which no document holds; its name still + # says what claims it. resolved, found = resolve( {"name": "gzip", "configuration": {"level": math.nan}}, CodecDefinition, SCOPE ) - assert (resolved.json, resolved.resolution) == (None, "invalid") + assert resolved == RefusedField( + json=UNSET, name="gzip", read_as=CodecDefinition, definition=GZIP_CODEC + ) assert _locs(found) == [(("configuration", "level"), "invalid_value")] def test_error_a_value_that_is_not_a_field() -> None: resolved, found = resolve(5, CodecDefinition, SCOPE) - assert resolved.resolution == "invalid" + assert resolved == RefusedField(json=5, name=None, read_as=CodecDefinition) assert len(found) == 1 @@ -515,7 +764,7 @@ def test_error_a_regular_grid_extent_is_negative() -> None: resolved, found = resolve( {"name": "regular", "configuration": {"chunk_shape": [2, -1]}}, ChunkGridDefinition, SCOPE ) - assert resolved.resolution == "invalid" + assert isinstance(resolved, RefusedField) assert _locs(found) == [(("configuration", "chunk_shape", 1), "invalid_value")] @@ -527,8 +776,8 @@ def test_error_a_nested_field_is_judged_where_it_sits() -> None: "configuration": {"codecs": ["crc32c", {"name": "gzip", "configuration": {"level": 12}}]}, } resolved, found = resolve(field, CodecDefinition, SCOPE) - assert resolved.resolution == "read" - assert resolved.nested[("codecs", 1)].resolution == "invalid" + assert isinstance(resolved, AcceptedField) + assert isinstance(resolved.nested[("codecs", 1)], RefusedField) assert _locs(found) == [ (("configuration", "codecs", 1, "configuration", "level"), "invalid_value") ] @@ -540,8 +789,8 @@ def test_error_a_nested_field_in_extra_items_is_judged_where_it_sits() -> None: "configuration": {"slow": {"name": "gzip", "configuration": {"level": 12}}}, } resolved, found = resolve(field, CodecDefinition, SCOPE) - assert resolved.resolution == "read" - assert resolved.nested[("slow",)].resolution == "invalid" + assert isinstance(resolved, AcceptedField) + assert isinstance(resolved.nested[("slow",)], RefusedField) assert _locs(found) == [(("configuration", "slow", "configuration", "level"), "invalid_value")] @@ -552,7 +801,7 @@ def test_error_a_nested_envelope_s_problem_is_its_own_and_the_rules_are_asked() resolved, found = resolve( {"name": "acme.stack", "configuration": {"codecs": [unread]}}, CodecDefinition, SCOPE ) - assert resolved.resolution == "read" + assert isinstance(resolved, AcceptedField) assert _locs(found) == [(("configuration", "codecs", 0, "must_understand"), "invalid_value")] # Its rules are asked all the same: one refuses the array -> bytes # codec it holds, which is the stack's own problem. @@ -561,7 +810,7 @@ def test_error_a_nested_envelope_s_problem_is_its_own_and_the_rules_are_asked() CodecDefinition, SCOPE, ) - assert resolved.resolution == "invalid" + assert isinstance(resolved, RefusedField) assert _locs(found) == [ (("configuration", "codecs", 1), "invalid_value"), (("configuration", "codecs", 0, "must_understand"), "invalid_value"), @@ -570,18 +819,18 @@ def test_error_a_nested_envelope_s_problem_is_its_own_and_the_rules_are_asked() def test_error_a_container_rule_is_not_asked_of_a_malformed_nested_field() -> None: # The stack's rule reads each nested field's name; one with no name is - # reported where it sits, and the rule is not asked, as `judge` would not. + # reported where it sits, and the rule is not asked, as `read_configuration` would not. field = {"name": "acme.stack", "configuration": {"codecs": [{"configuration": {}}]}} resolved, found = resolve(field, CodecDefinition, SCOPE) - assert resolved.resolution == "invalid" - assert _locs(found) == [(("configuration", "codecs", 0, "name"), "invalid_type")] - assert ACME_STACK.judge(field["configuration"])[0] is None + assert isinstance(resolved, RefusedField) + assert _locs(found) == [(("configuration", "codecs", 0, "name"), "missing_key")] + assert ACME_STACK.read_configuration(field["configuration"])[0] is None def test_error_a_rule_about_the_whole_configuration_lands_on_it() -> None: # A rule reports relative to the configuration: an empty location is # the configuration, judged alone or read in a field. - _, judged = ACME_PAIRED.judge({"first": 1}) + _, judged = ACME_PAIRED.read_configuration({"first": 1}) _, read = resolve( {"name": "acme.paired", "configuration": {"first": 1}}, CodecDefinition, SCOPE ) @@ -591,20 +840,20 @@ def test_error_a_rule_about_the_whole_configuration_lands_on_it() -> None: def test_error_a_rule_reads_the_fields_the_configuration_holds_as_the_scope_read_them() -> None: # `bytes` is an array -> bytes codec, which the stack's rule refuses - # from what the scope read; `judge` reads in no scope, so its rule sees + # from what the scope read; `read_configuration` reads in no scope, so its rule sees # nothing read. field = {"name": "acme.stack", "configuration": {"codecs": ["crc32c", "bytes"]}} resolved, found = resolve(field, CodecDefinition, SCOPE) - assert resolved.resolution == "invalid" + assert isinstance(resolved, RefusedField) assert _locs(found) == [(("configuration", "codecs", 1), "invalid_value")] - assert ACME_STACK.judge(field["configuration"])[1] == () + assert ACME_STACK.read_configuration(field["configuration"])[1] == () def test_error_a_nested_member_is_not_a_field() -> None: resolved, found = resolve( {"name": "acme.stack", "configuration": {"codecs": [5]}}, CodecDefinition, SCOPE ) - assert resolved.resolution == "invalid" + assert isinstance(resolved, RefusedField) assert _locs(found) == [(("configuration", "codecs", 0), "invalid_type")] @@ -620,17 +869,26 @@ def test_check_needs_nothing_but_the_value_and_a_typeddict() -> None: def test_check_reads_a_whole_array_document() -> None: - # A null in `dimension_names`, and a top-level key the document does - # not declare, typed by its `extra_items`. + # A null in `dimension_names`, a top-level key the document does not + # declare, typed by its `extra_items`, and a dimension of length 0. document = { - **ZarrV3ArrayMetadata.create_default(shape=(4,)).to_json(), - "dimension_names": [None], + **ZarrV3ArrayMetadata.create_default(shape=(0, 4)).to_json(), + "dimension_names": [None, "x"], "acme": {"must_understand": False}, } typed, problems = check(document, ZarrV3ArrayMetadataJSON) assert problems == () assert typed is not None - assert typed.get("dimension_names") == (None,) + assert typed.get("dimension_names") == (None, "x") + + +def test_error_check_holds_an_array_document_s_shape_to_its_bound() -> None: + document = {**ZarrV3ArrayMetadata.create_default(shape=(4,)).to_json(), "shape": [-1]} + typed, problems = check(document, ZarrV3ArrayMetadataJSON) + assert typed is None + assert [(found.loc, found.kind, dict(found.ctx)) for found in problems] == [ + (("shape", 0), "invalid_value", {"ge": 0}) + ] def test_configuration_of_types_what_its_definition_read() -> None: @@ -653,13 +911,33 @@ def test_error_check_judges_a_nested_envelope() -> None: assert [found.loc for found in problems] == [("codecs", 0, "configuration")] -def test_judge_is_the_check_and_then_the_rules() -> None: +def test_read_configuration_leaves_out_a_member_a_nested_field_s_envelope_does_not_declare() -> ( + None +): + # Reported as an unknown key and left out, as the checker leaves out a + # key a closed TypedDict does not declare: the configuration comes back. + typed, problems = ACME_STACK.read_configuration({"codecs": [{"name": "crc32c", "x": 1}]}) + assert typed == {"codecs": ({"name": "crc32c"},)} + assert _locs(problems) == [(("codecs", 0, "x"), "unknown_key")] + + +def test_error_read_configuration_refuses_a_nested_field_that_need_not_be_understood() -> None: + # A `must_understand` of false is no unknown key: the configuration + # does not come back. + typed, problems = ACME_STACK.read_configuration( + {"codecs": [{"name": "crc32c", "must_understand": False}]} + ) + assert typed is None + assert _locs(problems) == [(("codecs", 0, "must_understand"), "invalid_value")] + + +def test_read_configuration_is_the_check_and_then_the_rules() -> None: # A caller holding one configuration: the rules are asked only of a # configuration that type-checked, so they never meet a wrong type. - assert GZIP_CODEC.judge({"level": 5}) == ({"level": 5}, ()) - refused, problems = GZIP_CODEC.judge({"level": 12}) + assert GZIP_CODEC.read_configuration({"level": 5}) == ({"level": 5}, ()) + refused, problems = GZIP_CODEC.read_configuration({"level": 12}) assert (refused, _locs(problems)) == (None, [(("level",), "invalid_value")]) - mistyped, problems = GZIP_CODEC.judge({"level": "x"}) + mistyped, problems = GZIP_CODEC.read_configuration({"level": "x"}) assert (mistyped, _locs(problems)) == (None, [(("level",), "invalid_type")]) @@ -681,6 +959,174 @@ def test_a_scope_takes_a_name_over() -> None: assert scope.claimant(ChunkGridDefinition, "regular") is REGULAR_CHUNK_GRID +def test_a_scope_is_shown_by_how_many_definitions_it_holds() -> None: + # Short, as a validator's default argument is shown by `help`. + assert repr(CORE) == f"Context(<{len(CORE.definitions())} definitions>)" + + +def test_a_scope_pickles_as_its_definitions_and_copies_as_itself() -> None: + # A model holds the scope it was read in, and goes to another process with it. + again = pickle.loads(pickle.dumps(CORE_AND_EXTENSIONS)) + assert again.definitions() == CORE_AND_EXTENSIONS.definitions() + assert copy.copy(CORE) is CORE + assert copy.deepcopy(CORE) is CORE + + +LE: Final = {"name": "bytes", "configuration": {"endian": "little"}} +SHARD: Final = {"chunk_shape": [2], "codecs": [LE], "index_location": "end"} +NOSHUFFLE: Final = {"cname": "lz4", "clevel": 5, "shuffle": "noshuffle", "blocksize": 0} + + +@pytest.mark.parametrize( + ("one", "other", "kind", "equal"), + [ + ("uint8", {"name": "uint8"}, DataTypeDefinition, True), + ({"name": "bytes"}, {"name": "bytes", "configuration": {}}, CodecDefinition, True), + ( + {"name": "gzip", "configuration": {"level": 1}}, + {"name": "gzip", "configuration": {"level": 1}, "must_understand": True}, + CodecDefinition, + True, + ), + ("acme.codec", {"name": "acme.codec", "configuration": {}}, CodecDefinition, True), + ( + { + "name": "sharding_indexed", + "configuration": {**SHARD, "index_codecs": [LE, "crc32c"]}, + }, + { + "name": "sharding_indexed", + "configuration": {**SHARD, "index_codecs": [LE, {"name": "crc32c"}]}, + }, + CodecDefinition, + True, + ), + ( + {"name": "gzip", "configuration": {"level": 1}}, + {"name": "gzip", "configuration": {"level": 2}}, + CodecDefinition, + False, + ), + ( + {"name": "acme.codec", "configuration": {"a": 1}}, + {"name": "acme.codec", "configuration": {"a": 2}}, + CodecDefinition, + False, + ), + ("int8", "uint8", DataTypeDefinition, False), + # The spec's equivalences, which each definition's `canonical` + # folds: a member at its default, a run-length encoding, leading + # zeros in a size. + ( + {"name": "blosc", "configuration": {**NOSHUFFLE}}, + {"name": "blosc", "configuration": {**NOSHUFFLE, "typesize": 4}}, + CodecDefinition, + True, + ), + ( + "default", + {"name": "default", "configuration": {"separator": "/"}}, + ChunkKeyEncodingDefinition, + True, + ), + ( + "default", + {"name": "default", "configuration": {"separator": "."}}, + ChunkKeyEncodingDefinition, + False, + ), + ( + {"name": "v2"}, + {"name": "v2", "configuration": {"separator": "."}}, + ChunkKeyEncodingDefinition, + True, + ), + ( + { + "name": "sharding_indexed", + "configuration": {**SHARD, "index_codecs": [LE, "crc32c"]}, + }, + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [2], + "codecs": [LE], + "index_codecs": [LE, "crc32c"], + }, + }, + CodecDefinition, + True, + ), + ( + {"name": "zstd", "configuration": {"level": 3}}, + {"name": "zstd", "configuration": {"level": 3, "checksum": False}}, + CodecDefinition, + True, + ), + ( + {"name": "zstd", "configuration": {"level": 3}}, + {"name": "zstd", "configuration": {"level": 3, "checksum": True}}, + CodecDefinition, + False, + ), + ( + {"name": "rectilinear", "configuration": {"kind": "inline", "chunk_shapes": [[2, 2]]}}, + { + "name": "rectilinear", + "configuration": {"kind": "inline", "chunk_shapes": [[[2, 2]]]}, + }, + ChunkGridDefinition, + True, + ), + ("r008", "r8", DataTypeDefinition, True), + ], + ids=[ + "bare-or-object", + "no-or-empty-configuration", + "must-understand-true-or-absent", + "unclaimed-bare-or-object", + "a-field-it-holds-bare-or-object", + "another-configuration", + "unclaimed-another-configuration", + "another-name", + "blosc-typesize-noshuffle", + "default-separator", + "default-other-separator", + "v2-separator", + "sharding-index-location", + "zstd-checksum-false", + "zstd-checksum-true", + "rectilinear-rle", + "raw-bits-leading-zeros", + ], +) +def test_two_fields_are_equal_when_they_read_the_same( + one: JSONValue, other: JSONValue, kind: type[Definition[Any]], equal: bool +) -> None: + """However each was spelled: equal fields have one simplest spelling and one hash, though each is written as read.""" + first, _ = resolve(one, kind, CORE_AND_EXTENSIONS) + second, _ = resolve(other, kind, CORE_AND_EXTENSIONS) + assert isinstance(first, (AcceptedField, UnclaimedField)) + assert isinstance(second, (AcceptedField, UnclaimedField)) + assert (first == second) is equal + assert (canonical_of(first, ()) == canonical_of(second, ())) is equal + if equal: + assert hash(first) == hash(second) + + +def test_a_field_copied_or_pickled_is_read_by_a_definition_equal_to_its_own() -> None: + field, _ = resolve({"name": "gzip", "configuration": {"level": 1}}, CodecDefinition, CORE) + for again in (pickle.loads(pickle.dumps(field)), copy.deepcopy(field)): + assert again == field + assert configuration_of(again, GZIP_CODEC) == {"level": 1} + + +def test_a_definition_is_shown_by_its_kind_and_name() -> None: + # Short, as a reading that holds it shows it. + assert repr(GZIP_CODEC) == "CodecDefinition(name='gzip')" + assert repr(INT8_DATA_TYPE) == "DataTypeDefinition(name='int8')" + + class Unreadable(TypedDict, closed=True): members: set[int] @@ -694,7 +1140,7 @@ class AcmePlainConfiguration(TypedDict, closed=True): class AcmeDecimal(TypedDict, closed=True): - value: Decimal # a name the type checker sees, and the running module does not + value: Decimal # noqa: F821 - a name the running module does not define # pyright: ignore[reportUndefinedVariable] class AcmeUnresolvedConfiguration(TypedDict, closed=True): @@ -746,6 +1192,31 @@ def test_error_a_configuration_whose_annotations_do_not_resolve() -> None: ) +class AcmeUncheckedConfiguration(TypedDict, closed=True): + digits: Annotated[str, Predicate(str.isdigit)] + + +def test_error_a_configuration_with_a_constraint_the_checker_does_not_read() -> None: + # A bound the checker did not hold a value to would say what is not so. + with pytest.raises( + TypeError, + match=r"AcmeUncheckedConfiguration.digits: Predicate\(str.isdigit\) is not a constraint", + ): + CodecDefinition( + name="acme.unchecked", + configuration=AcmeUncheckedConfiguration, + kind="bytes_bytes", + size="dynamic", + ) + + +def test_error_a_fill_value_with_a_constraint_its_type_cannot_take() -> None: + with pytest.raises(TypeError, match="fill_value: ge: a bound is on a number, and a string"): + DataTypeDefinition( + name="acme.bounded", configuration=EmptyConfiguration, fill_value=Annotated[str, Ge(0)] + ) + + def test_error_a_definition_name_is_a_string() -> None: with pytest.raises(TypeError, match="a definition's name is a string"): CodecDefinition(name=5, configuration=Empty, kind="bytes_bytes", size="dynamic") # pyright: ignore[reportArgumentType] @@ -767,6 +1238,12 @@ def test_error_a_data_type_fill_value_no_checker_reads() -> None: DataTypeDefinition(name="acme.set", configuration=Empty, fill_value=set[int]) +def test_error_a_data_type_fill_value_holding_a_metadata_field() -> None: + # A value of the data type, which no scope reads as a field. + with pytest.raises(TypeError, match="'acme.f': fill_value: CodecField holds a metadata field"): + DataTypeDefinition(name="acme.f", configuration=Empty, fill_value=tuple[CodecField, ...]) + + @pytest.mark.parametrize( ("kind", "member"), [ @@ -800,3 +1277,196 @@ def test_error_a_field_is_read_as_a_kind_of_metadata(kind: type[Definition[Any]] def test_error_a_scope_refuses_a_definition_of_no_kind() -> None: with pytest.raises(TypeError, match="a definition of no kind"): Context.of(Definition(name="acme.kindless", configuration=Empty)) + + +@pytest.mark.parametrize( + ("name", "valid"), + [ + ("zstd", True), + ("numcodecs.adler32", True), + ("vlen-utf8", True), + ("acme_x", True), + ("r16", True), + ("https://example.com/codec", True), + ("urn:acme:codec", True), + ("urn:acme:%C3%BC", True), + ("x:", False), + ("Int8:", False), + ("a: b", False), + ("urn:acme:codec\n", False), + ("urn:acme:ü", False), + # Whitespace to Python's `re` but not ECMA-262, and the other way + # round: neither a URI character, so both dialects refuse them. + ("urn:a\x1c", False), + ("urn:a", False), + ("", False), + (" ", False), + ("Int8", False), + ("9x", False), + ("int8 ", False), + ("int8\n", False), + ("foo/bar", False), + ("-int8", False), + ("a", False), + ("r*", False), + ], +) +def test_an_extension_is_named_as_the_spec_names_one(name: str, valid: bool) -> None: + """The spec's regex `^[a-z][a-z0-9-_.]+$`, or a URI, which earlier versions of the spec required; anything else is refused before any definition is asked, by the field validator as by the reader.""" + for field, at in ((name, ()), ({"name": name}, ("name",))): + resolved, problems = resolve(field, CodecDefinition, CORE) + if valid: + assert not isinstance(resolved, RefusedField) + assert problems == () + else: + assert isinstance(resolved, RefusedField) + assert [(p.loc, p.kind) for p in problems] == [(at, "invalid_value")] + assert "expected an extension name" in problems[0].message + assert [(p.loc, p.kind) for p in validate_metadata_field_v3(field)] == ( + [] if valid else [(at, "invalid_value")] + ) + assert is_metadata_field_v3(field) is valid + + +def test_error_a_bad_name_is_a_problem_when_the_field_is_not_json_too() -> None: + # Refused before a definition is asked on either path: a field whose + # configuration is not JSON reports its name as the JSON path does, + # and no definition is asked to claim it. + field = {"name": "Acme", "configuration": {"x": float("nan")}} + resolved, problems = resolve(field, CodecDefinition, CORE) + assert isinstance(resolved, RefusedField) + assert resolved.definition is None + assert [(p.loc, p.kind) for p in problems] == [ + (("name",), "invalid_value"), + (("configuration", "x"), "invalid_value"), + ] + + +@pytest.mark.parametrize("name", ["Acme", "acme/x", "", "x"]) +def test_error_a_definition_is_named_as_the_spec_names_an_extension(name: str) -> None: + # No document could name it, so nothing would ever read with it. + with pytest.raises(TypeError, match="so no document names it"): + CodecDefinition( + name=name, configuration=EmptyConfiguration, kind="bytes_bytes", size="dynamic" + ) + + +def test_error_a_field_built_by_hand_is_named_as_the_spec_names_one() -> None: + with pytest.raises(TypeError, match="named as a document names a"): + UnclaimedField(json="Int8", name="Int8", read_as=CodecDefinition) + + +def test_error_a_field_without_a_name_is_missing_one() -> None: + resolved, problems = resolve({"configuration": {}}, CodecDefinition, CORE) + assert isinstance(resolved, RefusedField) + assert [(p.loc, p.kind, p.message) for p in problems] == [ + (("name",), "missing_key", "missing required key") + ] + + +@pytest.mark.parametrize("digits", [101, 4301]) +def test_raw_bits_of_more_than_a_hundred_digits_are_no_size_but_a_name(digits: int) -> None: + # `int` refuses to convert more than 4,300 digits, and no size has a + # hundred: such a name is an extension's, which nothing in scope claims. + resolved, problems = resolve("r" + "1" * digits, DataTypeDefinition, CORE_AND_EXTENSIONS) + assert type(resolved) is UnclaimedField + assert problems == () + + +def test_error_a_rule_that_writes_to_the_configuration_fails_there() -> None: + # A definition's functions are handed a read-only view of the field's + # own configuration, so no field holds what a function wrote. + def writes( + configuration: GzipCodecConfiguration, nested: Nested + ) -> Iterator[ValidationProblem]: + cast("dict[str, object]", configuration)["level"] = -1 + yield from () + + scope = CORE.extended_with(dataclasses.replace(GZIP_CODEC, rules=writes)) + with pytest.raises(TypeError, match="does not support item assignment") as raised: + resolve({"name": "gzip", "configuration": {"level": 1}}, CodecDefinition, scope) + assert raised.value.__notes__ == ["raised by the rules of 'gzip', reading ('configuration',)"] + + +def test_error_a_function_that_writes_to_the_configuration_fails_on_every_path() -> None: + # `read_configuration` hands the rules the same view `resolve` does, and `==` and + # `hash` hand `canonical` one, as `canonical_of` does. + def writes( + configuration: GzipCodecConfiguration, nested: Nested + ) -> Iterator[ValidationProblem]: + cast("dict[str, object]", configuration)["level"] = -1 + yield from () + + with pytest.raises(TypeError, match="does not support item assignment"): + dataclasses.replace(GZIP_CODEC, rules=writes).read_configuration({"level": 1}) + + def folds_in_place(configuration: GzipCodecConfiguration) -> GzipCodecConfiguration: + cast("dict[str, object]", configuration)["level"] = 0 + return configuration + + scope = CORE.extended_with(dataclasses.replace(GZIP_CODEC, canonical=folds_in_place)) + read, _ = resolve({"name": "gzip", "configuration": {"level": 1}}, CodecDefinition, scope) + with pytest.raises(TypeError, match="does not support item assignment"): + hash(read) + + +def test_error_a_canonical_that_raises_says_which_definition_raised_it() -> None: + def refuses(configuration: GzipCodecConfiguration) -> GzipCodecConfiguration: + raise ValueError("no") + + scope = CORE.extended_with(dataclasses.replace(GZIP_CODEC, canonical=refuses)) + read, _ = resolve({"name": "gzip", "configuration": {"level": 1}}, CodecDefinition, scope) + with pytest.raises(ValueError, match="no") as raised: + canonical_of(read, ()) + assert raised.value.__notes__ == ["raised by the canonical of 'gzip'"] + + +def test_error_a_field_s_configuration_and_nested_fields_are_read_only() -> None: + """What a field hands out -- its `json`, `configuration` and `nested` fields -- is read-only at every level, so a field cannot be put in a state its key, `==` and `refines` disagree about.""" + whole = {**SHARD, "index_codecs": [LE, {"name": "crc32c"}]} + shard, _ = resolve({"name": "sharding_indexed", "configuration": whole}, CodecDefinition, CORE) + assert isinstance(shard, AcceptedField) + with pytest.raises(TypeError): + shard.configuration["index_location"] = "start" # pyright: ignore[reportIndexIssue] + codecs = shard.configuration["codecs"] + assert isinstance(codecs, tuple) + inner = codecs[0] + assert isinstance(inner, Mapping) + with pytest.raises(TypeError): + inner["name"] = "crc32c" # pyright: ignore[reportIndexIssue] + with pytest.raises(TypeError): + del shard.nested[("codecs", 0)] # pyright: ignore[reportIndexIssue] + assert isinstance(shard.json, Mapping) + with pytest.raises(TypeError): + shard.json["name"] = "gzip" # pyright: ignore[reportIndexIssue] + unclaimed, _ = resolve({"name": "acme.x", "configuration": {"a": [1]}}, CodecDefinition, CORE) + assert isinstance(unclaimed, UnclaimedField) + with pytest.raises(TypeError): + unclaimed.configuration["a"] = 2 # pyright: ignore[reportIndexIssue] + refused, _ = resolve( + {"name": "sharding_indexed", "configuration": {**whole, "index_location": "x"}}, + CodecDefinition, + CORE, + ) + assert isinstance(refused, RefusedField) + with pytest.raises(TypeError): + del refused.nested[("codecs", 0)] # pyright: ignore[reportIndexIssue] + + +def test_a_field_holding_fields_is_copied_and_pickled_whole() -> None: + """A field pickles and deep-copies with the fields it holds, equal to itself, whether accepted, refused or left unclaimed.""" + whole = {**SHARD, "index_codecs": [LE, {"name": "crc32c"}]} + fields = [ + resolve({"name": "sharding_indexed", "configuration": whole}, CodecDefinition, CORE)[0], + resolve({"name": "acme.x", "configuration": {"a": [1]}}, CodecDefinition, CORE)[0], + resolve( + {"name": "sharding_indexed", "configuration": {**whole, "index_location": "x"}}, + CodecDefinition, + CORE, + )[0], + ] + for field in fields: + for again in (pickle.loads(pickle.dumps(field)), copy.deepcopy(field)): + assert again == field + assert type(again) is type(field) + assert again.nested == field.nested diff --git a/packages/zarr-metadata/tests/v3/test_every_definition.py b/packages/zarr-metadata/tests/v3/test_every_definition.py index 9df43313f0..2ea52fcc7f 100644 --- a/packages/zarr-metadata/tests/v3/test_every_definition.py +++ b/packages/zarr-metadata/tests/v3/test_every_definition.py @@ -12,20 +12,26 @@ import pytest +from zarr_metadata._json import value_at +from zarr_metadata.v3._definition import ( + canonical_of, + canonicalize, + configuration_of, +) from zarr_metadata.v3.codec.gzip import GZIP_CODEC from zarr_metadata.v3.data_type.raw import RAW_BYTES_DATA_TYPE, RawBytesConfiguration from zarr_metadata.v3.definition import ( CORE, CORE_AND_EXTENSIONS, + AcceptedField, ChunkGridDefinition, ChunkKeyEncodingDefinition, CodecDefinition, Context, DataTypeDefinition, Definition, + RefusedField, ValidationProblem, - canonicalize, - configuration_of, fill_value_problems, resolve, ) @@ -139,9 +145,10 @@ CASES = [(key, field) for key, fields in EXAMPLES.items() for field in fields] -def _read(key: str, field: object) -> tuple[str, list[tuple[tuple[str | int, ...], str]]]: +def _read(key: str, field: object) -> tuple[type, list[tuple[tuple[str | int, ...], str]]]: + """What the scope made of `field` -- `AcceptedField`, `UnclaimedField` or `RefusedField` -- and where each problem is.""" resolved, problems = resolve(field, KINDS[key.split(":")[0]], CORE_AND_EXTENSIONS) - return resolved.resolution, [(found.loc, found.kind) for found in problems] + return type(resolved), [(found.loc, found.kind) for found in problems] def _problems(key: str, field: object) -> list[tuple[tuple[str | int, ...], str]]: @@ -168,14 +175,16 @@ def test_every_definition_in_scope_has_an_example() -> None: ("key", "field"), CASES, ids=[f"{k}:{i}" for i, (k, _) in enumerate(CASES)] ) def test_every_example_reads_and_its_simplest_spelling_is_stable(key: str, field: object) -> None: - # Read, with nothing wrong; and its simplest spelling reads the same, + # AcceptedField, with nothing wrong; and its simplest spelling reads the same, # and is its own simplest spelling. - assert _read(key, field) == ("read", []) + assert _read(key, field) == (AcceptedField, []) kind = KINDS[key.split(":")[0]] simplest, problems = canonicalize(field, kind, CORE_AND_EXTENSIONS) assert problems == () assert simplest is not None assert canonicalize(simplest, kind, CORE_AND_EXTENSIONS) == (simplest, ()) + # A field read already is spelled as its JSON is, with nothing read again. + assert canonical_of(*resolve(field, kind, CORE_AND_EXTENSIONS)) == simplest @pytest.mark.parametrize( @@ -203,8 +212,12 @@ def test_every_example_reads_and_its_simplest_spelling_is_stable(key: str, field }, }, ), - # Nothing configured: the bare name, and no `must_understand: true`. - ({"name": "crc32c", "configuration": {}, "must_understand": True}, "crc32c"), + # Nothing configured: the name alone, as an object, which every reader + # takes, and no `must_understand: true`. + ({"name": "crc32c", "configuration": {}, "must_understand": True}, {"name": "crc32c"}), + # But a data type with nothing configured is its bare name, as core + # data types are written. + ({"name": "int8", "configuration": {}, "must_understand": True}, "int8"), # Nested fields each in their own simplest spelling. ( { @@ -222,10 +235,10 @@ def test_every_example_reads_and_its_simplest_spelling_is_stable(key: str, field "name": "sharding_indexed", "configuration": { "chunk_shape": (2,), - "codecs": ("bytes",), + "codecs": ({"name": "bytes"},), "index_codecs": ( {"name": "bytes", "configuration": {"endian": "little"}}, - "crc32c", + {"name": "crc32c"}, ), }, }, @@ -248,12 +261,63 @@ def test_every_example_reads_and_its_simplest_spelling_is_stable(key: str, field "configuration": {"kind": "inline", "chunk_shapes": (((32, 3), 16), 8)}, }, ), + # A member at its spec default is left out: a key encoding's + # separator, a shard's index location, zstd's checksum. + ({"name": "default", "configuration": {"separator": "/"}}, {"name": "default"}), + ({"name": "v2", "configuration": {"separator": "."}}, {"name": "v2"}), + ( + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [2], + "codecs": ["bytes"], + "index_codecs": [ + {"name": "bytes", "configuration": {"endian": "little"}}, + "crc32c", + ], + "index_location": "end", + }, + }, + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": (2,), + "codecs": ({"name": "bytes"},), + "index_codecs": ( + {"name": "bytes", "configuration": {"endian": "little"}}, + {"name": "crc32c"}, + ), + }, + }, + ), + ( + {"name": "zstd", "configuration": {"level": 3, "checksum": False}}, + {"name": "zstd", "configuration": {"level": 3}}, + ), + ], + ids=[ + "blosc-noshuffle", + "name-alone", + "data-type-bare-name", + "sharding-nested", + "cast-value-target", + "rectilinear-rle", + "default-separator", + "v2-separator", + "sharding-index-location", + "zstd-checksum", ], - ids=["blosc-noshuffle", "bare-name", "sharding-nested", "cast-value-target", "rectilinear-rle"], ) def test_the_simplest_spelling(field: dict[str, Any], simplest: object) -> None: - kind = ChunkGridDefinition if field["name"] == "rectilinear" else CodecDefinition + kinds = { + "rectilinear": ChunkGridDefinition, + "int8": DataTypeDefinition, + "default": ChunkKeyEncodingDefinition, + "v2": ChunkKeyEncodingDefinition, + } + kind = kinds.get(field["name"], CodecDefinition) assert canonicalize(field, kind, CORE_AND_EXTENSIONS) == (simplest, ()) + assert canonical_of(*resolve(field, kind, CORE_AND_EXTENSIONS)) == simplest @pytest.mark.parametrize( @@ -271,7 +335,8 @@ def test_raw_bits_read_as_r_star_with_the_size_their_name_carries( field: object, bits: int, simplest: str ) -> None: resolved, problems = resolve(field, DataTypeDefinition, CORE_AND_EXTENSIONS) - assert (resolved.resolution, problems) == ("read", ()) + assert problems == () + assert isinstance(resolved, AcceptedField) assert resolved.definition is RAW_BYTES_DATA_TYPE assert resolved.json == field assert configuration_of(resolved, RAW_BYTES_DATA_TYPE) == {"bits": bits} @@ -284,22 +349,58 @@ def test_a_reader_reads_raw_bits_its_own_way_by_defining_r_star() -> None: mine = DataTypeDefinition(name="r*", configuration=RawBytesConfiguration) scope = CORE_AND_EXTENSIONS.extended_with(mine) resolved, problems = resolve("r12", DataTypeDefinition, scope) + assert problems == () + assert isinstance(resolved, AcceptedField) assert resolved.definition is mine - assert (resolved.resolution, problems) == ("read", ()) - again = scope.extended_with(RAW_BYTES_DATA_TYPE) - assert resolve("r12", DataTypeDefinition, again)[0].definition is RAW_BYTES_DATA_TYPE + # The package's own takes no 12 bits, and refuses them. + again, _ = resolve("r12", DataTypeDefinition, scope.extended_with(RAW_BYTES_DATA_TYPE)) + assert isinstance(again, RefusedField) + assert again.definition is RAW_BYTES_DATA_TYPE -@pytest.mark.parametrize("field", ["r*", {"name": "r*", "configuration": {"bits": 16}}]) -def test_r_star_is_notation_that_names_nothing(field: object) -> None: +@pytest.mark.parametrize( + ("field", "at"), + [("r*", ()), ({"name": "r*", "configuration": {"bits": 16}}, ("name",))], + ids=["bare", "object"], +) +def test_error_r_star_is_notation_and_no_name(field: object, at: tuple[str, ...]) -> None: # How the specification's table writes raw bits, and no document's name - # for them: read as any name nothing in scope claims. + # for them: `*` is no character of an extension name. resolved, problems = resolve(field, DataTypeDefinition, CORE_AND_EXTENSIONS) - assert (resolved.resolution, resolved.definition, problems) == ("out_of_scope", None, ()) + assert type(resolved) is RefusedField + assert [(p.loc, p.kind) for p in problems] == [(at, "invalid_value")] + + +@pytest.mark.parametrize( + ("field", "kind", "simplest"), + [ + ({"name": "zfpy"}, CodecDefinition, {"name": "zfpy"}), + # What it simplifies to is its own definition's call, so its + # configuration is kept as written; its envelope is its kind's. + ( + {"name": "zfpy", "configuration": {"level": [1]}}, + CodecDefinition, + {"name": "zfpy", "configuration": {"level": (1,)}}, + ), + ("zfpy", CodecDefinition, {"name": "zfpy"}), + ( + {"name": "zfpy", "configuration": {}, "must_understand": True}, + CodecDefinition, + {"name": "zfpy"}, + ), + ({"name": "acme.decimal"}, DataTypeDefinition, "acme.decimal"), + ], + ids=["object", "configured", "bare-codec", "spelled-out", "data-type"], +) +def test_an_unclaimed_field_keeps_its_configuration_in_its_kind_s_envelope( + field: object, kind: type[Definition[Any]], simplest: object +) -> None: + assert canonicalize(field, kind, CORE) == (simplest, ()) -def test_an_unclaimed_field_keeps_its_own_spelling() -> None: - assert canonicalize({"name": "zfpy"}, CodecDefinition, CORE) == ({"name": "zfpy"}, ()) +def _shard(codecs: list[object], index_codecs: list[object]) -> dict[str, object]: + configuration = {"chunk_shape": [2], "codecs": codecs, "index_codecs": index_codecs} + return {"name": "sharding_indexed", "configuration": configuration} @pytest.mark.parametrize( @@ -331,9 +432,19 @@ def test_an_unclaimed_field_keeps_its_own_spelling() -> None: ( {"name": "zfpy", "configuration": {}, "extra": 1}, CodecDefinition, - [(("extra",), "invalid_value")], + [(("extra",), "unknown_key")], ), (None, CodecDefinition, [((), "invalid_type")]), + ( + _shard(["bytes", {"name": "gzip", "configuration": {"level": 12}}], ["bytes"]), + CodecDefinition, + [(("configuration", "codecs", 1, "configuration", "level"), "invalid_value")], + ), + ( + _shard(["bytes"], ["bytes", {"name": "gzip", "configuration": {"level": 1}}]), + CodecDefinition, + [(("configuration", "index_codecs", 1), "invalid_value")], + ), ], ids=[ "refused-value", @@ -342,16 +453,21 @@ def test_an_unclaimed_field_keeps_its_own_spelling() -> None: "must-understand-false", "stray", "null", + "holding-a-refused-field", + "a-codec-of-dynamic-size-where-one-of-static-size-goes", ], ) def test_error_a_field_with_a_problem_has_no_simplest_spelling( field: dict[str, Any] | None, kind: type[Definition[Any]], found: list[object] ) -> None: # Whatever the author wrote stays theirs: a simpler spelling would - # drop the unknown key, the stray member or the `must_understand`. + # drop the unknown key, the stray member or the `must_understand`, or + # spell what does not hold. A field read already, given its problems, + # has none either. simplest, problems = canonicalize(field, kind, CORE_AND_EXTENSIONS) assert simplest is None assert [(problem.loc, problem.kind) for problem in problems] == found + assert canonical_of(*resolve(field, kind, CORE_AND_EXTENSIONS)) is None def test_error_a_canonical_that_does_not_hold_is_the_definitions_fault() -> None: @@ -392,6 +508,17 @@ def test_error_blosc_typesize_is_missing_while_shuffling() -> None: assert _one("codecs:blosc", configuration) == [(("configuration", "typesize"), "missing_key")] +def test_a_rule_is_asked_only_of_a_configuration_within_its_bounds() -> None: + # As pydantic's after-validators and zod's refinements are: a rule + # relies on the bounds its type declares, so a value out of them is + # the one problem reported until it is fixed. + shuffled = {key: value for key, value in BLOSC.items() if key != "typesize"} + assert _one("codecs:blosc", {**shuffled, "clevel": 12}) == [ + (("configuration", "clevel"), "invalid_value") + ] + assert _one("codecs:blosc", shuffled) == [(("configuration", "typesize"), "missing_key")] + + def test_error_blosc_typesize_is_not_positive() -> None: assert _one("codecs:blosc", {**BLOSC, "typesize": 0}) == [ (("configuration", "typesize"), "invalid_value") @@ -414,6 +541,13 @@ def test_error_scale_offset_scalar_is_null() -> None: ] +def test_error_regular_chunk_length_is_zero() -> None: + # Along a dimension of length 0 too: a grid's chunks have a size. + assert _one("chunk_grid:regular", {"chunk_shape": [4, 0]}) == [ + (("configuration", "chunk_shape", 1), "invalid_value") + ] + + def test_error_sharding_inner_chunk_extent_is_zero() -> None: configuration = {"chunk_shape": [0], "codecs": ["bytes"], "index_codecs": ["bytes"]} assert _one("codecs:sharding_indexed", configuration) == [ @@ -421,6 +555,100 @@ def test_error_sharding_inner_chunk_extent_is_zero() -> None: ] +@pytest.mark.parametrize( + ("kind", "field", "loc", "ctx"), + [ + ( + CodecDefinition, + {"name": "gzip", "configuration": {"level": 10}}, + ("configuration", "level"), + {"ge": 0, "le": 9}, + ), + ( + CodecDefinition, + {"name": "zstd", "configuration": {"level": -131073}}, + ("configuration", "level"), + {"ge": -131072, "le": 22}, + ), + ( + CodecDefinition, + {"name": "blosc", "configuration": {**BLOSC, "clevel": -1}}, + ("configuration", "clevel"), + {"ge": 0, "le": 9}, + ), + ( + CodecDefinition, + {"name": "blosc", "configuration": {**BLOSC, "blocksize": -1}}, + ("configuration", "blocksize"), + {"ge": 0}, + ), + ( + CodecDefinition, + { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [2, 0], + "codecs": ["bytes"], + "index_codecs": ["bytes"], + }, + }, + ("configuration", "chunk_shape", 1), + {"ge": 1}, + ), + ( + ChunkGridDefinition, + {"name": "regular", "configuration": {"chunk_shape": [2, 0]}}, + ("configuration", "chunk_shape", 1), + {"ge": 1}, + ), + ( + ChunkGridDefinition, + {"name": "rectilinear", "configuration": {"kind": "inline", "chunk_shapes": [0]}}, + ("configuration", "chunk_shapes", 0), + {"ge": 1}, + ), + ( + ChunkGridDefinition, + { + "name": "rectilinear", + "configuration": {"kind": "inline", "chunk_shapes": [[4, [2, 0]]]}, + }, + ("configuration", "chunk_shapes", 0, 1, 1), + {"ge": 1}, + ), + ( + DataTypeDefinition, + {"name": "numpy.timedelta64", "configuration": {"unit": "s", "scale_factor": 2**31}}, + ("configuration", "scale_factor"), + {"ge": 1, "le": 2**31 - 1}, + ), + ], + ids=[ + "gzip-level", + "zstd-level", + "blosc-clevel", + "blosc-blocksize", + "sharding-inner-chunk-extent", + "regular-chunk-extent", + "rectilinear-extent", + "rectilinear-run-count", + "numpy-time-scale-factor", + ], +) +def test_a_bound_is_its_member_s_type_and_its_problem_holds_it( + kind: type[Definition[Any]], + field: dict[str, Any], + loc: tuple[str | int, ...], + ctx: dict[str, int], +) -> None: + # Declared on the TypedDict, as pydantic reads a bound, and not in a + # rule: the problem holds the bound, and what was found. + _, problems = resolve(field, kind, CORE_AND_EXTENSIONS) + assert [(p.loc, p.kind, p.input, dict(p.ctx)) for p in problems] == [ + (loc, "invalid_value", value_at(field, loc), ctx) + ] + + STATIC_SIZE = ("bytes", "cast_value", "crc32c", "scale_offset", "transpose") DYNAMIC_SIZE = ("blosc", "gzip", "sharding_indexed", "zstd") @@ -635,7 +863,7 @@ def test_error_a_configuration_written_beside_a_raw_bits_name() -> None: # The name carries the configuration, so one written beside it holds # nothing: each member is a key nothing declares. assert _read("data_type:r*", {"name": "r16", "configuration": {"bits": 16}}) == ( - "read", + AcceptedField, [(("configuration", "bits"), "unknown_key")], ) @@ -651,7 +879,7 @@ def test_a_rule_is_a_function_over_the_typeddict() -> None: # configuration can ask them without a scope or a field around it. from zarr_metadata.v3.codec.blosc import BLOSC_CODEC - _, problems = BLOSC_CODEC.judge({**BLOSC, "clevel": 10}) + _, problems = BLOSC_CODEC.read_configuration({**BLOSC, "clevel": 10}) assert problems == ( ValidationProblem(("clevel",), "expected an integer in [0, 9], got 10", "invalid_value"), ) diff --git a/packages/zarr-metadata/tests/v3/test_fill_values.py b/packages/zarr-metadata/tests/v3/test_fill_values.py index 10be671772..f64aacf4f9 100644 --- a/packages/zarr-metadata/tests/v3/test_fill_values.py +++ b/packages/zarr-metadata/tests/v3/test_fill_values.py @@ -8,23 +8,34 @@ from __future__ import annotations +import dataclasses +import json import math -from typing import TYPE_CHECKING +from typing import TYPE_CHECKING, Any, cast, get_args import pytest +from hypothesis import given +from hypothesis import strategies as st from typing_extensions import TypedDict -from zarr_metadata.model import validate_array_metadata_v3, validate_group_metadata_v3 +from zarr_metadata._json import JSON_DEPTH, value_at +from zarr_metadata._sentinel import UNSET +from zarr_metadata.model import ( + validate_array_metadata_v3, + validate_group_metadata_v3, +) from zarr_metadata.model._array import ZarrV3ArrayMetadata +from zarr_metadata.v3.data_type._float import FloatWidth, complex_fill_value_rules, float_bits from zarr_metadata.v3.data_type.struct import STRUCT_DATA_TYPE from zarr_metadata.v3.definition import ( CORE_AND_EXTENSIONS, + AcceptedField, DataTypeDefinition, EmptyConfiguration, JSONValue, Nested, - Resolved, ValidationProblem, + canonical_fill_value, fill_value_problems, resolve, ) @@ -202,6 +213,31 @@ def test_error_a_byte_value_out_of_range( assert _problems(data_type, fill_value) == [(loc, "invalid_value")] +@pytest.mark.parametrize( + ("data_type", "fill_value", "loc", "ctx"), + [ + ("int8", 128, (), {"ge": -128, "le": 127}), + ("uint16", -1, (), {"ge": 0, "le": 2**16 - 1}), + ("int64", 2**63, (), {"ge": -(2**63), "le": 2**63 - 1}), + ("uint64", 2**64, (), {"ge": 0, "le": 2**64 - 1}), + ("r16", [1, 256], (1,), {"ge": 0, "le": 255}), + ("bytes", [-1], (0,), {"ge": 0, "le": 255}), + (DATETIME, 2**63, (), {"ge": -(2**63), "le": 2**63 - 1}), + ], + ids=["int8", "uint16", "int64", "uint64", "raw-bits-byte", "bytes-byte", "numpy-time"], +) +def test_a_fill_value_s_range_is_its_type_s( + data_type: JSONValue, fill_value: object, loc: tuple[int, ...], ctx: dict[str, int] +) -> None: + # `Int8FillValue` is `Annotated[int, Interval(ge=-128, le=127)]`: the + # problem holds the range, and what was found. + resolved, _ = resolve(data_type, DataTypeDefinition, CORE_AND_EXTENSIONS) + problems = fill_value_problems(resolved, fill_value) + assert [(p.loc, p.input, dict(p.ctx)) for p in problems] == [ + (loc, value_at(fill_value, loc), ctx) + ] + + @pytest.mark.parametrize("fill_value", ["!!", "AQI"]) def test_error_a_bytes_fill_value_that_is_not_base64(fill_value: str) -> None: assert _problems("bytes", fill_value) == [((), "invalid_value")] @@ -265,7 +301,7 @@ def test_error_a_key_the_fill_value_shape_does_not_declare_hides_no_rule() -> No ] -def test_a_fill_value_nested_hundreds_deep_is_read() -> None: +def test_a_fill_value_nested_as_deep_as_a_reader_walks_is_read() -> None: def deep(levels: int) -> dict[str, object]: value: dict[str, object] = {} for _ in range(levels): @@ -275,7 +311,8 @@ def deep(levels: int) -> dict[str, object]: document = dict(ZarrV3ArrayMetadata.create_default().to_json()) | { "data_type": STRUCT, "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}], - "fill_value": {"a": 1, "b": 0.5, "c": deep(600)}, + # `c` sits two levels down, and holds the deepest value the cap admits. + "fill_value": {"a": 1, "b": 0.5, "c": deep(JSON_DEPTH - 3)}, } assert [(p.loc, p.kind) for p in validate_array_metadata_v3(document)] == [ (("fill_value", "c"), "unknown_key") @@ -285,5 +322,204 @@ def deep(levels: int) -> dict[str, object]: def test_a_struct_read_without_its_field_types_leaves_its_fields_unjudged() -> None: # A reading built by hand, holding no field type's reading. configuration = {"fields": ({"name": "a", "data_type": "int8"},)} - struct = Resolved(STRUCT, "read", STRUCT_DATA_TYPE, configuration) + struct = AcceptedField( + json=STRUCT, name="struct", definition=STRUCT_DATA_TYPE, configuration=configuration + ) assert fill_value_problems(struct, {"a": 300}) == () + + +# --- canonical spellings --------------------------------------------------- + + +def _alike(left: object, right: object) -> bool: + """Whether two JSON values are written alike: `==` takes `-0.0` for `0.0`.""" + return json.dumps(left, sort_keys=True) == json.dumps(right, sort_keys=True) + + +def _read(data_type: JSONValue) -> AcceptedField[DataTypeDefinition[Any]]: + resolved, found = resolve(data_type, DataTypeDefinition, CORE_AND_EXTENSIONS) + assert found == () + assert isinstance(resolved, AcceptedField) + return resolved + + +@pytest.mark.parametrize( + ("data_type", "spellings", "canonical"), + [ + ("float32", ["NaN", "0x7fc00000", "0x7FC00000"], "NaN"), + # Any other NaN is its bits: the spec names the one. + ("float32", ["0xffc00000"], "0xffc00000"), + ("float32", ["0x7fc00001", "0x7FC00001"], "0x7fc00001"), + ("float32", ["Infinity", "0x7f800000", 3.5e38, 10**39], "Infinity"), + ("float32", ["-Infinity", "0xff800000", -(10**39)], "-Infinity"), + ("float32", [0, 0.0, "0x00000000", 1e-46], 0.0), + # Zero's sign is a value of its own. + ("float32", [-0.0, "0x80000000"], -0.0), + ("float32", [1, 1.0, "0x3f800000"], 1.0), + # The number of the fewest digits that rounds to the value. + ("float32", [0.1, 0.10000000149011612, "0x3dcccccd"], 0.1), + # A number is read as a float64, as a JSON parser reads one, and + # then rounded to the type, as zarrs, tensorstore and numpy round + # it: an integer and a number with a fraction that read as one + # float64 are one value, and an integer whose float64 is halfway + # between two float32 values is the even one, as they store it. + ("float32", [2**53 + 2**29 + 1, 9007199791611905.0, 2**53], 9007199000000000.0), + ("float32", [2**60 + 2**36 + 1, 2**60, 1.1529215e18], 1.1529215e18), + # Of the numbers of the fewest digits, the one past the nearest: at + # a power of two, those that round to it reach further above it. + ("float16", [0.015625, "0x2400"], 0.01563), + ("float32", ["0x6b000000"], 1.5474251e26), + ("float16", [65504, 65519, "0x7bff"], 65500.0), + ("float16", [65520, "Infinity", "0x7c00"], "Infinity"), + # Halfway rounds to the even value. + ("float16", [2049, 2048], 2048.0), + ("float16", [2051, 2052], 2052.0), + ("float16", ["NaN", "0x7e00", "0x7E00"], "NaN"), + ("float16", [-0.0, "0x8000"], -0.0), + ("float64", [2**53 + 1, 2**53], float(2**53)), + ("float64", [10**400, "Infinity", "0x7ff0000000000000"], "Infinity"), + ("float64", [-(10**400), "-Infinity", "0xfff0000000000000"], "-Infinity"), + ("float64", ["NaN", "0x7ff8000000000000"], "NaN"), + ("complex64", [["NaN", -0.0], ["0x7fc00000", "0x80000000"]], ("NaN", -0.0)), + ("complex128", [[0.1, 1], ["0x3fb999999999999a", 1.0]], (0.1, 1.0)), + ("bytes", [[65], "QQ==", "QR=="], "QQ=="), + ("bytes", [[], ""], ""), + (DATETIME, ["NaT", -(2**63)], "NaT"), + (TIMEDELTA, ["NaT", -(2**63)], "NaT"), + (TIMEDELTA, [5], 5), + (STRUCT, [{"a": 1, "b": "0x7fc00000"}, {"b": "NaN", "a": 1}], {"a": 1, "b": "NaN"}), + # A field type nothing in scope claims spells its fill value as written. + ( + { + "name": "struct", + "configuration": {"fields": [{"name": "a", "data_type": "acme.decimal"}]}, + }, + [{"a": -0.0}], + {"a": -0.0}, + ), + # A type that spells each value one way spells it as written. + ("int8", [-128], -128), + ("bool", [True], True), + ("string", ["NaN"], "NaN"), + ("r16", [[0, 1]], (0, 1)), + ], +) +def test_a_fill_value_s_canonical_spelling_is_its_value_s( + data_type: JSONValue, spellings: list[object], canonical: JSONValue +) -> None: + read = _read(data_type) + for spelling in spellings: + spelled = canonical_fill_value(read, spelling) + assert _alike(spelled, canonical), (spelling, spelled) + # A fill value of the type, spelled as itself. + assert fill_value_problems(read, canonical) == () + assert _alike(canonical_fill_value(read, canonical), canonical) + + +def test_a_data_type_nothing_in_scope_claims_spells_a_fill_value_as_written() -> None: + unclaimed, _ = resolve("acme.decimal", DataTypeDefinition, CORE_AND_EXTENSIONS) + values: list[JSONValue] = [-0.0, True, 1, [1.0], None] + for value in values: + assert _alike(canonical_fill_value(unclaimed, value), value) + + +WIDTHS: tuple[FloatWidth, ...] = get_args(FloatWidth) + + +@given(st.data()) +def test_a_float_s_canonical_spelling_spells_its_bits(data: st.DataObject) -> None: + # Every value of each float type, NaNs among them: its canonical + # spelling is a fill value of the type, spelling the same bits, and + # its own canonical spelling. + width = data.draw(st.sampled_from(WIDTHS)) + # Drawn by part, so normal values, infinities and NaNs each turn up: + # bits drawn whole are almost all subnormals. + fraction = {16: 10, 32: 23, 64: 52}[width] + sign = data.draw(st.integers(0, 1)) + exponent = data.draw(st.integers(0, 2 ** (width - 1 - fraction) - 1)) + mantissa = data.draw(st.integers(0, 2**fraction - 1)) + bits = (sign << (width - 1)) | (exponent << fraction) | mantissa + read = _read(f"float{width}") + spelled = canonical_fill_value(read, f"0x{bits:0{width // 4}x}") + assert spelled is not UNSET + assert fill_value_problems(read, spelled) == () + assert float_bits(cast("float | str", spelled), width) == bits + assert _alike(canonical_fill_value(read, spelled), spelled) + + +def test_error_a_fill_value_with_a_problem_has_no_canonical_spelling() -> None: + # `UNSET`, since `None` is the JSON null, a fill value of a data type + # the scope did not read. + assert canonical_fill_value(_read("float32"), "0x7fc0") is UNSET + assert canonical_fill_value(_read("int8"), 1.0) is UNSET + assert canonical_fill_value(_read("int8"), math.nan) is UNSET + unclaimed, _ = resolve("acme.decimal", DataTypeDefinition, CORE_AND_EXTENSIONS) + assert canonical_fill_value(unclaimed, None) is None + + +def test_error_a_canonical_spelling_that_raises_says_which_data_type_raised_it() -> None: + def refuses(configuration: EmptyConfiguration, nested: Nested, value: object) -> JSONValue: + raise ValueError("no") + + scope = CORE_AND_EXTENSIONS.extended_with( + dataclasses.replace(ACME_POINT, fill_value_canonical=refuses) + ) + resolved, _ = resolve("acme.point", DataTypeDefinition, scope) + with pytest.raises(ValueError, match="no") as raised: + canonical_fill_value(resolved, {"x": 1}) + assert raised.value.__notes__ == ["raised by the fill_value_canonical of 'acme.point'"] + + +def test_error_a_canonical_spelling_that_is_not_json_is_refused() -> None: + def not_json(configuration: EmptyConfiguration, nested: Nested, value: object) -> JSONValue: + return cast("JSONValue", {1, 2}) + + scope = CORE_AND_EXTENSIONS.extended_with( + dataclasses.replace(ACME_POINT, fill_value_canonical=not_json) + ) + resolved, _ = resolve("acme.point", DataTypeDefinition, scope) + with pytest.raises(TypeError, match="'acme.point': its fill_value_canonical gives JSON"): + canonical_fill_value(resolved, {"x": 1}) + + +@pytest.mark.parametrize( + ("name", "low", "high"), + [ + ("int8", -(2**7), 2**7 - 1), + ("int16", -(2**15), 2**15 - 1), + ("int32", -(2**31), 2**31 - 1), + ("int64", -(2**63), 2**63 - 1), + ("uint8", 0, 2**8 - 1), + ("uint16", 0, 2**16 - 1), + ("uint32", 0, 2**32 - 1), + ("uint64", 0, 2**64 - 1), + ], +) +def test_an_integer_type_takes_exactly_its_range(name: str, low: int, high: int) -> None: + read = _read(name) + assert fill_value_problems(read, low) == () + assert fill_value_problems(read, high) == () + for outside in (low - 1, high + 1): + (problem,) = fill_value_problems(read, outside) + assert (problem.loc, problem.kind, dict(problem.ctx)) == ( + (), + "invalid_value", + {"ge": low, "le": high}, + ) + + +def test_a_complex_fill_value_keeps_what_its_component_rules_found() -> None: + """A complex fill value's problems are its component rules' own, each moved under its component's index, with the `input` and `ctx` they carry.""" + + def at_least_one( + configuration: EmptyConfiguration, nested: Nested, part: object + ) -> Iterator[ValidationProblem]: + if part == 0: + yield ValidationProblem( + ("x",), "expected >= 1", "invalid_value", input=0, ctx={"ge": 1} + ) + + rules = complex_fill_value_rules(at_least_one) + (found,) = rules({}, {}, (1.0, 0)) + assert (found.loc, found.message, found.kind) == ((1, "x"), "expected >= 1", "invalid_value") + assert (found.input, dict(found.ctx)) == (0, {"ge": 1}) diff --git a/packages/zarr-metadata/tests/v3/test_grid_shapes.py b/packages/zarr-metadata/tests/v3/test_grid_shapes.py index 6797fcb47e..00e0cf753e 100644 --- a/packages/zarr-metadata/tests/v3/test_grid_shapes.py +++ b/packages/zarr-metadata/tests/v3/test_grid_shapes.py @@ -11,13 +11,18 @@ import pytest -from zarr_metadata.model import ZarrV3ArrayMetadata, validate_array_metadata_v3 +from zarr_metadata.model import ( + ZarrV3ArrayMetadata, + validate_array_metadata_v3, +) +from zarr_metadata.v3._definition import ( + chunk_grid_lengths, +) from zarr_metadata.v3.codec.crc32c import Empty from zarr_metadata.v3.definition import ( CORE_AND_EXTENSIONS, ChunkGridDefinition, JSONValue, - chunk_grid_lengths, resolve, ) @@ -44,9 +49,9 @@ def _problems(grid: JSONValue, shape: tuple[int, ...]) -> list[tuple[tuple[str | [ (_regular(), (), ()), (_regular(4, 4), (10, 3), ({4}, {4})), - # A chunk longer than its dimension, and a chunk length of 0 for a - # dimension of length 0. - (_regular(8, 0), (3, 0), ({8}, {0})), + # A chunk longer than its dimension, and a chunk over a dimension of + # length 0. + (_regular(8, 1), (3, 0), ({8}, {1})), # A bare integer repeats until it covers its dimension. (_rectilinear(4), (10,), ({4},)), (_rectilinear([4, 4, 2]), (10,), ({4, 2},)), @@ -79,12 +84,6 @@ def test_error_a_regular_grid_of_another_rank(grid: JSONValue, shape: tuple[int, assert _problems(grid, shape) == [(("configuration", "chunk_shape"), "invalid_value")] -def test_error_a_regular_chunk_length_of_0_for_a_dimension_that_is_not_empty() -> None: - assert _problems(_regular(4, 0), (4, 3)) == [ - (("configuration", "chunk_shape", 1), "invalid_value") - ] - - @pytest.mark.parametrize( ("grid", "shape"), [(_rectilinear(4), (10, 3)), (_rectilinear(4, 4), (1,))] ) diff --git a/packages/zarr-metadata/tests/v3/test_hierarchy.py b/packages/zarr-metadata/tests/v3/test_hierarchy.py new file mode 100644 index 0000000000..4b9944bf2b --- /dev/null +++ b/packages/zarr-metadata/tests/v3/test_hierarchy.py @@ -0,0 +1,256 @@ +"""A Zarr v3 hierarchy: the names and paths of its nodes, and the tree they make, as the core specification constrains them. + +Names: +https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L818-L837 +Paths: +https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L211-L229 +The tree: +https://github.com/zarr-developers/zarr-specs/blob/fc7dd9c9beb5a50b87f9b08b00bf50fc0048482f/docs/v3/core/index.rst#L177-L181 +""" + +from __future__ import annotations + +import time +from typing import TYPE_CHECKING, Literal, cast + +import pytest + +from zarr_metadata.model import ( + MetadataValidationError, + ValidationProblem, + is_node_name_v3, + is_node_path_v3, + parse_node_name_v3, + parse_node_path_v3, + validate_node_name_v3, + validate_node_path_v3, +) +from zarr_metadata.v3._hierarchy import hierarchy_problems + +if TYPE_CHECKING: + from collections.abc import Callable + + +@pytest.mark.parametrize( + ("value", "message"), + [ + # The root's name, and names of any code points but the reserved. + ("", None), + ("a", None), + ("FOO", None), + ("a.b", None), + ("...a", None), + ("_a", None), + ("a__", None), + ("zarr.json.bak", None), + ("ñ", None), + ("/", 'expected a node name, got "/", which holds "/"'), + ("a/b", 'expected a node name, got "a/b", which holds "/"'), + (".", 'expected a node name, got ".", which is periods alone'), + ("..", 'expected a node name, got "..", which is periods alone'), + ("__a", 'expected a node name, got "__a", which starts with the reserved "__"'), + ("zarr.json", 'expected a node name, got "zarr.json", which is the reserved "zarr.json"'), + # Every reason a name is not one, in one problem. + ( + "__/", + 'expected a node name, got "__/", which holds "/" and starts with the reserved "__"', + ), + ], +) +def test_node_names(value: str, message: str | None) -> None: + problems = validate_node_name_v3(value) + assert [problem.message for problem in problems] == ([] if message is None else [message]) + assert {(problem.loc, problem.kind, problem.input) for problem in problems} <= { + ((), "invalid_value", value) + } + assert is_node_name_v3(value) is (message is None) + if message is None: + assert parse_node_name_v3(value) == value + + +@pytest.mark.parametrize( + ("value", "message"), + [ + # The root's path, and the paths below it. + ("/", None), + ("/a", None), + ("/a/b", None), + ("/a/.b/c..", None), + ("", 'expected a node path, got "", which does not start with "/"'), + ("a/b", 'expected a node path, got "a/b", which does not start with "/"'), + ("/a/", 'expected a node path, got "/a/", which ends with "/"'), + ( + "//", + 'expected a node path, got "//", which ends with "/" and holds an empty name between two "/"', + ), + ("/a//b", 'expected a node path, got "/a//b", which holds an empty name between two "/"'), + ( + "/a/../b", + 'expected a node path, got "/a/../b", which holds "..", a name that is periods alone', + ), + ( + "/__a", + ( + 'expected a node path, got "/__a", which holds "__a", a name that starts with the ' + 'reserved "__"' + ), + ), + ( + "/a/zarr.json", + ( + 'expected a node path, got "/a/zarr.json", which holds "zarr.json", a name that is ' + 'the reserved "zarr.json"' + ), + ), + # Every reason a path is not one, in one problem: the first name that + # is not a node name is said, and the rest counted. + ( + "/./__a/", + ( + 'expected a node path, got "/./__a/", which ends with "/", holds ".", a name that is ' + "periods alone and holds 1 more name that is not a node name" + ), + ), + ], +) +def test_node_paths(value: str, message: str | None) -> None: + problems = validate_node_path_v3(value) + assert [problem.message for problem in problems] == ([] if message is None else [message]) + assert {(problem.loc, problem.kind, problem.input) for problem in problems} <= { + ((), "invalid_value", value) + } + assert is_node_path_v3(value) is (message is None) + if message is None: + assert parse_node_path_v3(value) == value + + +def test_error_parse_node_name_raises_every_reason_a_string_is_not_a_node_name() -> None: + with pytest.raises(MetadataValidationError) as raised: + parse_node_name_v3("__/") + assert raised.value.problems == validate_node_name_v3("__/") + assert len(raised.value.problems) == 1 + assert 'holds "/"' in raised.value.problems[0].message + assert 'starts with the reserved "__"' in raised.value.problems[0].message + + +def test_error_parse_node_path_raises_every_reason_a_string_is_not_a_node_path() -> None: + with pytest.raises(MetadataValidationError) as raised: + parse_node_path_v3("/./__a/") + assert raised.value.problems == validate_node_path_v3("/./__a/") + assert len(raised.value.problems) == 1 + assert 'ends with "/"' in raised.value.problems[0].message + assert "1 more name" in raised.value.problems[0].message + + +@pytest.mark.parametrize( + ("validate", "parse", "what"), + [ + (validate_node_name_v3, parse_node_name_v3, "a node name"), + (validate_node_path_v3, parse_node_path_v3, "a node path"), + ], + ids=["name", "path"], +) +def test_error_a_node_name_or_path_is_a_string( + validate: Callable[[object], tuple[ValidationProblem, ...]], + parse: Callable[[object], str], + what: str, +) -> None: + problems = validate(3) + assert problems == (ValidationProblem((), f"expected {what}, got 3", "invalid_type"),) + assert problems[0].input == 3 + with pytest.raises(MetadataValidationError) as raised: + parse(b"a") + assert [(problem.loc, problem.kind) for problem in raised.value.problems] == [ + ((), "invalid_type") + ] + + +@pytest.mark.parametrize( + "nodes", + [ + # No nodes, a root alone, and a tree of groups whose leaves are + # arrays; the root may be an array, which is then the hierarchy. + {}, + {"/": "group"}, + {"/": "array"}, + {"/": "group", "/a": "array", "/g": "group", "/g/b": "array", "/g/h": "group"}, + # A node of no type known is taken as a group. + {"/": "group", "/x": None, "/x/a": "array"}, + ], + ids=["empty", "root-group", "root-array", "tree", "unknown-type"], +) +def test_a_hierarchy(nodes: dict[str, Literal["array", "group"] | None]) -> None: + assert hierarchy_problems(nodes) == () + + +def test_error_a_hierarchy_holds_its_nodes_at_node_paths() -> None: + assert hierarchy_problems({"/": "group", "a": "array", "/b/": "array"}) == ( + ValidationProblem( + ("a",), 'expected a node path, got "a", which does not start with "/"', "invalid_value" + ), + ValidationProblem( + ("/b/",), 'expected a node path, got "/b/", which ends with "/"', "invalid_value" + ), + ) + + +@pytest.mark.parametrize( + ("nodes", "path"), + [ + ({"/": "group", "/a": "array", "/a/b": "group"}, "/a/b"), + # However many groups are missing between them: none would help. + ({"/": "group", "/a": "array", "/a/b/c": "array"}, "/a/b/c"), + # The root is an array: the hierarchy is that array alone. + ({"/": "array", "/a": "group"}, "/a"), + ], + ids=["child", "descendant", "below-the-root"], +) +def test_error_no_node_is_below_an_array( + nodes: dict[str, Literal["array", "group"] | None], path: str +) -> None: + holder = next(ancestor for ancestor in ("/a", "/") if nodes.get(ancestor) == "array") + message = f'expected a node below a group, got "{path}", below the array "{holder}"' + assert hierarchy_problems(nodes) == (ValidationProblem((path,), message, "invalid_value"),) + + +def test_error_a_hierarchy_holds_the_group_holding_each_node() -> None: + # The nearest group missing above a node, once, counting those above it: + # the root's among them. + assert hierarchy_problems({"/a/b/c": "array", "/a/b/d": "array"}) == ( + ValidationProblem( + ("/a/b",), 'missing the group holding "/a/b/c", and 2 groups above it', "missing_key" + ), + ) + assert hierarchy_problems({"/": "group", "/a/b": "array"}) == ( + ValidationProblem(("/a",), 'missing the group holding "/a/b"', "missing_key"), + ) + + +def test_problems_stay_proportional_to_the_path() -> None: + # One problem per value, however many names are wrong and however many + # groups are missing, so a hostile path costs what it weighs. + path = "/" * 10_000 + problems = validate_node_path_v3(path) + assert len(problems) == 1 + assert len(problems[0].message) < len(path) + 200 + deep = "/a" * 5_000 + found = hierarchy_problems({deep: "array"}) + assert [(problem.loc, problem.kind) for problem in found] == [ + ((deep.rpartition("/")[0],), "missing_key") + ] + assert found[0].message.endswith(", and 4999 groups above it") + assert len(found[0].message) < len(deep) + 200 + + +def test_a_hierarchy_is_walked_in_time_proportional_to_its_paths() -> None: + # Each path split once and walked name by name: a key of a million + # characters, and thousands of keys under one long missing prefix, + # each cost what they weigh, where walking up by ancestor strings + # costs the square. + started = time.perf_counter() + long = "/a" * 500_000 + assert len(hierarchy_problems({long: "array"})) == 1 + prefix = "/a" * 5_000 + many = {f"{prefix}/{index}": "array" for index in range(2_000)} + assert len(hierarchy_problems(cast("dict[str, Literal['array', 'group'] | None]", many))) == 1 + assert time.perf_counter() - started < 10 diff --git a/packages/zarr-metadata/tests/v3/test_kinds.py b/packages/zarr-metadata/tests/v3/test_kinds.py new file mode 100644 index 0000000000..af05932bed --- /dev/null +++ b/packages/zarr-metadata/tests/v3/test_kinds.py @@ -0,0 +1,262 @@ +"""Kinds of definition: the class declared with `kind=True`, open to kinds of another format.""" + +from __future__ import annotations + +from collections.abc import Mapping +from dataclasses import dataclass +from typing import TYPE_CHECKING, Annotated, Any, ClassVar, TypeAlias, TypeVar, cast + +import pytest +from annotated_types import Ge +from typing_extensions import TypedDict + +from zarr_metadata._json import ValidationProblem +from zarr_metadata.v3._definition import ( + DataTypeDefinition, + WithFillValue, + as_kind, + field_json_schema, + kind_of, +) +from zarr_metadata.v3._scope import kind_name +from zarr_metadata.v3.codec.gzip import GZIP_CODEC +from zarr_metadata.v3.definition import ( + AcceptedField, + CodecDefinition, + Context, + Definition, + EmptyConfiguration, + RefusedField, + UnclaimedField, + canonical_fill_value, + fill_value_problems, + resolve, +) + +if TYPE_CHECKING: + from zarr_metadata._common import JSONValue + +C = TypeVar("C") + + +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class MyCodec(CodecDefinition[Any]): + """A codec definition with a member of its own: still a codec.""" + + note: str = "" + + +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class Tag(Definition[C], kind=True, format=3): + """A kind of its own, filed apart from every v3 kind.""" + + label: ClassVar[str] = "tag" + + +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class NoKind(Definition[Any]): + """A definition subclass that declares no kind.""" + + +def test_a_subclass_of_a_kind_is_a_definition_of_that_kind() -> None: + """A definition built as a subclass of `CodecDefinition` is filed, read and named as a codec: the kind is the nearest class in the MRO declared with `kind=True`, not the class of the definition.""" + mine = MyCodec(name="mine", configuration=EmptyConfiguration, kind="bytes_bytes", size="static") + assert kind_of(mine) is CodecDefinition + scope = Context.of(mine) + assert scope.claimant(CodecDefinition, "mine") is mine + assert isinstance(resolve({"name": "mine"}, CodecDefinition, scope)[0], AcceptedField) + + +def test_a_kind_of_its_own_is_filed_apart() -> None: + """A class declared with `kind=True` is a kind: `as_kind` accepts it with or without type arguments, a scope files its definitions apart from every other kind's, scopes compare by what each files, and messages name the kind by its label.""" + tag = Tag(name="tag1", configuration=EmptyConfiguration) + scope = Context.of(tag, GZIP_CODEC) + assert as_kind(Tag) is Tag + assert as_kind(Tag[Any]) is Tag + assert scope.claimant(Tag, "tag1") is tag + assert scope.claimant(CodecDefinition, "tag1") is None + assert scope == Context.of(GZIP_CODEC, tag) + assert kind_name(Tag) == "tag" + assert kind_name(CodecDefinition) == "codec" + + +def test_error_a_class_that_declares_no_kind_is_of_none() -> None: + """A `Definition` subclass declared without `kind=True` is of no kind: `kind_of` is None, `Context.of` refuses a definition of it, and `as_kind` refuses the class.""" + none = NoKind(name="nokind", configuration=EmptyConfiguration) + assert kind_of(none) is None + with pytest.raises(TypeError, match="a definition of no kind"): + Context.of(none) + with pytest.raises(TypeError, match="is not a kind of metadata"): + as_kind(NoKind) + + +class Params(TypedDict, closed=True): + level: Annotated[int, Ge(0)] + + +Loc = tuple[str | int, ...] + + +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class Flat(Definition[C], kind=True, format=3): + """A kind whose format writes the parameters beside the name: `{"id": name, **parameters}`.""" + + label: ClassVar[str] = "flat" + + @classmethod + def named_configuration( + cls, value: object + ) -> tuple[str | None, Mapping[str, object] | None, tuple[ValidationProblem, ...]]: + if not isinstance(value, Mapping): + return None, None, () + entry = cast("Mapping[str, object]", value) + name = entry.get("id") + if not isinstance(name, str): + return None, None, () + return name, {key: item for key, item in entry.items() if key != "id"}, () + + @classmethod + def envelope_problems(cls, value: object) -> tuple[ValidationProblem, ...]: + if cls.named_configuration(value)[0] is None: + return (ValidationProblem((), "expected an object with a string 'id'", "invalid_type"),) + return () + + @classmethod + def envelope_json(cls, name: str, configuration: Mapping[str, JSONValue]) -> JSONValue: + return {"id": name, **configuration} + + @classmethod + def configuration_loc(cls, loc: Loc) -> Loc: + return loc + + @classmethod + def name_loc(cls, loc: Loc) -> Loc: + return (*loc, "id") + + +FLAT = Flat(name="flat", configuration=Params) +FLAT_SCOPE = Context.of(FLAT) + + +@pytest.mark.parametrize( + ("field", "kind", "problems", "written"), + [ + ({"id": "flat", "level": 1}, AcceptedField, [], {"id": "flat", "level": 1}), + ({"id": "other", "x": 1}, UnclaimedField, [], {"id": "other", "x": 1}), + ({"id": "flat", "level": -1}, RefusedField, [(("c", "level"), "invalid_value")], None), + ( + {"id": "flat", "payload": object()}, + RefusedField, + [(("c", "payload"), "invalid_type")], + None, + ), + ("flat", RefusedField, [(("c",), "invalid_type")], None), + ], + ids=["read", "unclaimed", "out-of-range", "not-json", "not-an-object"], +) +def test_a_kind_reads_the_envelope_its_format_writes( + field: object, kind: type, problems: list[tuple[Loc, str]], written: object +) -> None: + """`resolve` reads a field as its kind's classmethods say the format writes one: the name and parameters are split as the kind splits them, problems sit where the kind puts the configuration, and a field read or unclaimed is written back in the kind's envelope.""" + resolved, found = resolve(field, Flat, FLAT_SCOPE, ("c",)) + assert type(resolved) is kind + assert [(problem.loc, problem.kind) for problem in found] == problems + if written is not None: + assert not isinstance(resolved, RefusedField) + assert resolved.to_json() == written + + +@dataclass(frozen=True, kw_only=True, slots=True, repr=False) +class Typed(WithFillValue[C], kind=True, format=3): + """A kind of another format whose definitions take a fill value.""" + + label: ClassVar[str] = "typed" + + +Count: TypeAlias = Annotated[int, Ge(0)] | None + +TYPED = Typed(name="typed", configuration=EmptyConfiguration, fill_value=Count) + + +def test_a_fill_value_is_judged_by_any_kind_with_one() -> None: + """`fill_value_problems` and `canonical_fill_value` judge a fill value by the definition's `fill_value` members whatever kind it is, since the members live on `WithFillValue`, which every data type kind derives from.""" + resolved, _ = resolve("typed", Typed, Context.of(TYPED)) + assert fill_value_problems(resolved, 3) == () + assert fill_value_problems(resolved, None) == () + assert [(p.loc, p.kind) for p in fill_value_problems(resolved, -1, ("fill_value",))] == [ + (("fill_value",), "invalid_value") + ] + assert canonical_fill_value(resolved, 3) == 3 + + +def test_error_the_json_schema_writer_writes_v3_fields_only() -> None: + """`field_json_schema` writes the v3 envelope, so a kind of another format is refused with a `TypeError` naming the limit rather than a wrong schema.""" + from zarr_metadata.v2.definition import CORE_V2, ZarrV2DataTypeDefinition + + with pytest.raises(TypeError, match="Zarr v3"): + field_json_schema(ZarrV2DataTypeDefinition, CORE_V2) + with pytest.raises(TypeError, match="Zarr v3"): + field_json_schema(Tag, Context.of(Tag(name="tag1", configuration=EmptyConfiguration))) + + +def test_a_v3_value_that_is_no_field_is_shown_in_the_message() -> None: + """A nested value that is not a metadata field at all is reported with the value shown, as it was before the kind's envelope judged it.""" + _, problems = resolve( + {"name": "sharding_indexed", "configuration": {"chunk_shape": [1], "codecs": [3]}}, + CodecDefinition, + Context.of( + *__import__("zarr_metadata.v3.definition", fromlist=["CORE"]).CORE.definitions() + ), + ) + assert any(p.message.endswith("got 3") for p in problems), [p.message for p in problems] + + +def test_error_a_kind_declares_its_format() -> None: + """A kind says which Zarr format its fields belong to, `format=2` or `format=3`, in the class header beside `kind=True`: a kind declared without one, or with a format that is neither, is a `TypeError` at class creation.""" + with pytest.raises(TypeError, match="format"): + type("NoFormat", (Definition,), {}, kind=True) + with pytest.raises(TypeError, match="format"): + type("WrongFormat", (Definition,), {}, kind=True, format=4) + with pytest.raises(TypeError, match="format"): + type("FormatOnADefinition", (CodecDefinition,), {}, format=3) + + +TAG = Tag(name="tag", configuration=EmptyConfiguration) + + +def test_a_scope_has_the_format_of_the_kinds_it_files() -> None: + """A scope's `format` is the one format of every kind it files: 3 for `CORE`, 2 for `CORE_V2`, and None for a scope that files nothing, which reads in either format.""" + from zarr_metadata.v2.definition import CORE_V2 + from zarr_metadata.v3.definition import CORE + + assert CORE.format == 3 + assert CORE_V2.format == 2 + assert Context.of().format is None + assert CORE.extended_with(TAG).format == 3 + + +def test_error_a_scope_files_one_format() -> None: + """Definitions of two formats cannot share a scope: `Context.of`, `extended_with` and `joined` each raise `TypeError` naming both formats.""" + from zarr_metadata.v2.data_type.scalar import UINT_V2 + from zarr_metadata.v2.definition import CORE_V2 + from zarr_metadata.v3.definition import CORE + + with pytest.raises(TypeError, match="format"): + Context.of(TAG, UINT_V2) + with pytest.raises(TypeError, match="format"): + CORE.extended_with(UINT_V2) + with pytest.raises(TypeError, match="format"): + Context.joined(CORE, CORE_V2) + + +def test_error_a_field_is_read_in_a_scope_of_its_format() -> None: + """`resolve` refuses a scope of another format than the kind's with `TypeError`, and reads in a scope of no format, which claims nothing.""" + from zarr_metadata.v2.definition import CORE_V2, ZarrV2DataTypeDefinition + from zarr_metadata.v3.definition import CORE + + with pytest.raises(TypeError, match="format"): + resolve(" Chunk: @pytest.mark.parametrize( ("lengths", "data_type", "match"), - [((4, 2), None, "a chunk's lengths are"), (None, "float32", "a chunk's data type is")], + [ + ((4, 2), None, "a chunk's lengths are"), + (None, "float32", "a chunk's data type is"), + # A field read as a codec, though nothing claims it. + (None, resolve("acme.t", CodecDefinition, SCOPE)[0], "a chunk's data type is"), + ], ) def test_error_a_transition_that_builds_a_chunk_of_something_else( lengths: object, data_type: object, match: str @@ -446,8 +456,10 @@ class AcmeHolderConfiguration(TypedDict, closed=True): types: tuple[DataTypeField, ...] -def _holder(pipelines: object) -> Resolved[CodecDefinition[Any]]: - """A codec holding a pipeline of codecs and a list of data types, whose pipelines are `pipelines`.""" +def _holder( + pipelines: object, types: JSONValue = ("uint8",) +) -> ResolvedField[CodecDefinition[Any]]: + """A codec holding a pipeline of codecs and a list of data types, `types`, whose pipelines are `pipelines`.""" holder = CodecDefinition( name="acme.holder", configuration=AcmeHolderConfiguration, @@ -455,7 +467,7 @@ def _holder(pipelines: object) -> Resolved[CodecDefinition[Any]]: size="dynamic", pipelines=pipelines, # pyright: ignore[reportArgumentType] ) - field = {"name": "acme.holder", "configuration": {"codecs": [LITTLE], "types": ["uint8"]}} + field = {"name": "acme.holder", "configuration": {"codecs": [LITTLE], "types": types}} return resolve(field, CodecDefinition, SCOPE.extended_with(holder))[0] @@ -465,9 +477,19 @@ def test_error_pipelines_that_give_something_else(given: object) -> None: read_pipeline([_holder(lambda configuration, nested, chunk: given)], CHUNK) -@pytest.mark.parametrize("member", ["nowhere", "types"]) -def test_error_pipelines_that_name_a_member_holding_no_codecs(member: str) -> None: - holder = _holder(lambda configuration, nested, chunk: {member: Chunk()}) +@pytest.mark.parametrize( + ("member", "types"), + [ + ("nowhere", ("uint8",)), + ("types", ("uint8",)), + # Data types nothing in scope claims are still read as data types. + ("types", ("acme.t",)), + ], +) +def test_error_pipelines_that_name_a_member_holding_no_codecs( + member: str, types: JSONValue +) -> None: + holder = _holder(lambda configuration, nested, chunk: {member: Chunk()}, types) with pytest.raises(TypeError, match=f"its pipelines name {member!r}, which holds no list"): read_pipeline([holder], CHUNK) @@ -479,3 +501,29 @@ def test_error_pipelines_that_raise_say_whose_they_are() -> None: assert raised.value.__notes__ == [ "raised by the pipelines of 'acme.holder', reading ('codecs', 0, 'configuration')" ] + + +def test_error_a_stage_s_inner_pipelines_are_read_only() -> None: + """The pipelines a stage holds of a shard's inner codecs cannot be changed in place, and a reading holding them pickles and deep-copies equal to itself.""" + import copy + import pickle + + shard: JSONValue = { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [2], + "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}], + "index_codecs": [ + {"name": "bytes", "configuration": {"endian": "little"}}, + {"name": "crc32c"}, + ], + "index_location": "end", + }, + } + stages, _ = _read([shard], CHUNK) + stage = stages[0] + assert "codecs" in stage.inner + with pytest.raises(TypeError): + stage.inner["codecs"] = () # pyright: ignore[reportIndexIssue] + for again in (pickle.loads(pickle.dumps(stages)), copy.deepcopy(stages)): + assert again == stages diff --git a/packages/zarr-metadata/tests/v3/test_scope.py b/packages/zarr-metadata/tests/v3/test_scope.py new file mode 100644 index 0000000000..c0a0872fca --- /dev/null +++ b/packages/zarr-metadata/tests/v3/test_scope.py @@ -0,0 +1,482 @@ +"""The algebra of scopes: value semantics for `Context`, claims, refinement, disagreements and joins.""" + +from __future__ import annotations + +import pickle +from typing import Any, get_type_hints + +import pytest +from hypothesis import given +from hypothesis import strategies as st + +from zarr_metadata.v3._definition import ( + fields_of, +) +from zarr_metadata.v3._scope import ( + Claims, + Disagreements, + claims_of, + kind_name, + refines, +) +from zarr_metadata.v3.codec.bytes import BYTES_CODEC +from zarr_metadata.v3.codec.crc32c import CRC32C_CODEC, Empty +from zarr_metadata.v3.codec.gzip import GZIP_CODEC +from zarr_metadata.v3.codec.sharding_indexed import SHARDING_INDEXED_CODEC +from zarr_metadata.v3.codec.zstd import ZSTD_CODEC +from zarr_metadata.v3.data_type.bytes import BYTES_DATA_TYPE +from zarr_metadata.v3.data_type.raw import RAW_BYTES_DATA_TYPE +from zarr_metadata.v3.definition import ( + CORE, + CORE_AND_EXTENSIONS, + ChunkGridDefinition, + ChunkKeyEncodingDefinition, + CodecDefinition, + Conflict, + Context, + DataTypeDefinition, + Definition, + RefusedField, + ResolvedField, + ScopeConflictError, + StorageTransformerDefinition, + resolve, +) + +SHARD = { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [1], + "codecs": ["bytes"], + "index_codecs": [{"name": "bytes", "configuration": {"endian": "little"}}, "crc32c"], + }, +} + +MY_GZIP = CodecDefinition(name="gzip", configuration=Empty, kind="bytes_bytes", size="dynamic") +"""A private reading of the name `gzip`: another definition under one name, which takes no configuration. + +Definitions compare by what they hold, so one rebuilt from the core +TypedDict with the core rules would be the core definition; this one +reads `gzip` otherwise. +""" + + +def _read(data: object, kind: type[Definition[Any]], scope: Context) -> ResolvedField[Any]: + return resolve(data, kind, scope)[0] + + +@pytest.mark.parametrize( + ("left", "right", "equal"), + [ + (Context.of(GZIP_CODEC, BYTES_CODEC), Context.of(BYTES_CODEC, GZIP_CODEC), True), + (CORE, pickle.loads(pickle.dumps(CORE)), True), + (CORE, CORE_AND_EXTENSIONS, False), + (Context.of(GZIP_CODEC), Context.of(GZIP_CODEC, ZSTD_CODEC), False), + (Context.of(), Context.of(), True), + ], + ids=["order", "pickle", "core-vs-extensions", "subset", "empty"], +) +def test_a_scope_is_equal_to_another_by_the_definitions_it_files( + left: Context, right: Context, equal: bool +) -> None: + """Two scopes are one when they file the same definitions under the same names, however they were built; equal scopes hash alike.""" + assert (left == right) is equal + if equal: + assert hash(left) == hash(right) + + +def test_a_scope_is_a_set_member() -> None: + """A scope hashes, so it can key a dict or sit in a set, which it could not when its tables were mapping proxies.""" + assert len({CORE, CORE_AND_EXTENSIONS, pickle.loads(pickle.dumps(CORE))}) == 2 + + +def test_error_a_scope_conflict_says_each_disagreement() -> None: + """A scope conflict lists each `(kind, name)` with what was claimed and what was found, and where, so a caller can see every disagreement at once.""" + error = ScopeConflictError( + ( + Conflict((CodecDefinition, "bytes"), BYTES_CODEC, None, ("codecs", 0)), + Conflict((CodecDefinition, "gzip"), None, GZIP_CODEC), + ) + ) + assert error.conflicts[0].key == (CodecDefinition, "bytes") + assert str(error) == ( + "codec 'bytes' at ('codecs', 0): claimed CodecDefinition(name='bytes') of " + "BytesCodecConfiguration, found no definition; " + "codec 'gzip': claimed no definition, found CodecDefinition(name='gzip') of " + "GzipCodecConfiguration" + ) + + +@pytest.mark.parametrize( + ("field", "claims"), + [ + (_read("gzip", CodecDefinition, CORE), {(CodecDefinition, "gzip"): GZIP_CODEC}), + (_read("zstd", CodecDefinition, CORE), {(CodecDefinition, "zstd"): None}), + ( + _read("r16", DataTypeDefinition, CORE), + {(DataTypeDefinition, "r*"): RAW_BYTES_DATA_TYPE}, + ), + ( + _read(SHARD, CodecDefinition, CORE), + { + (CodecDefinition, "sharding_indexed"): SHARDING_INDEXED_CODEC, + (CodecDefinition, "bytes"): BYTES_CODEC, + (CodecDefinition, "crc32c"): CRC32C_CODEC, + }, + ), + ( + _read({"name": "gzip", "configuration": {"level": 12}}, CodecDefinition, CORE), + {(CodecDefinition, "gzip"): GZIP_CODEC}, + ), + (RefusedField(json=3, name=None, read_as=CodecDefinition), {}), + ], + ids=["read", "unclaimed", "raw-bits", "nested", "refused-claimed", "refused-nameless"], +) +def test_claims_of_says_what_a_reading_claimed_of_each_name( + field: ResolvedField[Any], claims: dict[object, object] +) -> None: + """A reading's claims name the definition that read each name the field and the fields it holds write, keyed as the scope files it -- raw bits under `r*` -- and None where nothing claimed one; a field refused by a definition still claims it, and one that names nothing claims nothing.""" + assert claims_of(fields_of(field)) == claims + + +def test_error_claims_of_refuses_one_name_read_two_ways() -> None: + """Fields read in two scopes that give one name two definitions have no single set of claims: a `ScopeConflictError` naming the key.""" + fields = [ + *fields_of(_read("gzip", CodecDefinition, CORE), ("codecs", 0)), + *fields_of(_read("gzip", CodecDefinition, Context.of(MY_GZIP)), ("codecs", 1)), + ] + with pytest.raises(ScopeConflictError) as raised: + claims_of(fields) + (conflict,) = raised.value.conflicts + assert (conflict.key, conflict.claimed, conflict.found, conflict.loc) == ( + (CodecDefinition, "gzip"), + GZIP_CODEC, + MY_GZIP, + ("codecs", 1), + ) + + +GZIP_FIELD = {"name": "gzip", "configuration": {"level": 5}} +ZSTD_FIELD = {"name": "zstd", "configuration": {"level": 3}} +BLOSC = {"cname": "zstd", "shuffle": "noshuffle", "blocksize": 0} +BLOSC_LEVEL_ONE = {"name": "blosc", "configuration": {**BLOSC, "clevel": 1}} +BLOSC_LEVEL_TRUE = {"name": "blosc", "configuration": {**BLOSC, "clevel": True}} +"""Two documents Python's `==` takes for one, which `json_text` tells apart.""" +SHARD_HOLDING_REFUSED = { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [1], + "codecs": ["bytes", {"name": "gzip", "configuration": {"level": 12}}], + "index_codecs": [{"name": "bytes", "configuration": {"endian": "little"}}, "crc32c"], + }, +} +"""A shard read, holding a gzip its definition refuses.""" +NESTED_ZSTD = { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [1], + "codecs": ["bytes", ZSTD_FIELD], + "index_codecs": [{"name": "bytes", "configuration": {"endian": "little"}}, "crc32c"], + }, +} + + +@pytest.mark.parametrize( + ("field", "other", "expected"), + [ + (_read(GZIP_FIELD, CodecDefinition, CORE), _read(GZIP_FIELD, CodecDefinition, CORE), True), + ( + _read({"name": "crc32c"}, CodecDefinition, CORE), + _read("crc32c", CodecDefinition, CORE), + True, + ), + ( + _read(ZSTD_FIELD, CodecDefinition, CORE_AND_EXTENSIONS), + _read(ZSTD_FIELD, CodecDefinition, CORE), + True, + ), + ( + _read(ZSTD_FIELD, CodecDefinition, CORE), + _read(ZSTD_FIELD, CodecDefinition, CORE_AND_EXTENSIONS), + False, + ), + ( + _read(GZIP_FIELD, CodecDefinition, CORE), + _read(GZIP_FIELD, CodecDefinition, Context.of(MY_GZIP)), + False, + ), + ( + _read(NESTED_ZSTD, CodecDefinition, CORE_AND_EXTENSIONS), + _read(NESTED_ZSTD, CodecDefinition, CORE), + True, + ), + ( + _read(NESTED_ZSTD, CodecDefinition, CORE), + _read(NESTED_ZSTD, CodecDefinition, CORE_AND_EXTENSIONS), + False, + ), + ( + _read({"name": "gzip", "configuration": {"level": 12}}, CodecDefinition, CORE), + _read("gzip", CodecDefinition, CORE), + False, + ), + (_read("zstd", CodecDefinition, CORE), _read("zstd", CodecDefinition, CORE), True), + ( + _read({"name": "zstd", "configuration": {"level": 1}}, CodecDefinition, CORE), + _read({"name": "zstd", "configuration": {"level": 2}}, CodecDefinition, CORE), + False, + ), + ( + _read("crc32c", CodecDefinition, CORE), + _read({"name": "crc32c"}, CodecDefinition, Context.of()), + True, + ), + ( + _read(BLOSC_LEVEL_ONE, CodecDefinition, CORE), + _read(BLOSC_LEVEL_TRUE, CodecDefinition, Context.of()), + False, + ), + ( + _read(SHARD_HOLDING_REFUSED, CodecDefinition, CORE), + _read(SHARD_HOLDING_REFUSED, CodecDefinition, CORE), + True, + ), + ], + ids=[ + "same", + "same-spelled-otherwise", + "gain", + "loss", + "conflict", + "nested-gain", + "nested-loss", + "refused", + "both-unclaimed", + "unclaimed-differ", + "gain-over-another-spelling", + "no-gain-over-true-for-one", + "holding-a-refused-field", + ], +) +def test_refines_orders_readings_by_information( + field: ResolvedField[Any], other: ResolvedField[Any], expected: bool +) -> None: + """`field` refines `other` when it reads the same where both read and gains where `other` left a name unclaimed -- in the fields it holds too; a loss, a conflict, a refused field, or two unclaimed fields written differently do not.""" + assert refines(field, other) is expected + + +DOCUMENTS = [ + GZIP_FIELD, + "crc32c", + {"name": "crc32c"}, + {"name": "crc32c", "configuration": {}}, + ZSTD_FIELD, + NESTED_ZSTD, + BLOSC_LEVEL_ONE, + SHARD_HOLDING_REFUSED, +] +SCOPES = [Context.of(), CORE, CORE_AND_EXTENSIONS] +READINGS = [_read(document, CodecDefinition, scope) for document in DOCUMENTS for scope in SCOPES] + + +@given(st.sampled_from(READINGS), st.sampled_from(READINGS), st.sampled_from(READINGS)) +def test_refines_is_a_partial_order_whose_bottom_is_equality( + a: ResolvedField[Any], b: ResolvedField[Any], c: ResolvedField[Any] +) -> None: + """Over readings of documents in several scopes and spellings, `refines` is reflexive and transitive, two fields that refine each other are equal, and equal fields refine the same fields.""" + assert refines(a, a) + if refines(a, b) and refines(b, c): + assert refines(a, c) + assert (refines(a, b) and refines(b, a)) is (a == b) + if a == b: + assert refines(c, a) is refines(c, b) + + +@pytest.mark.parametrize( + ("scope", "claims", "gains", "conflicts"), + [ + (CORE, {(CodecDefinition, "gzip"): GZIP_CODEC}, (), ()), + ( + CORE_AND_EXTENSIONS, + {(CodecDefinition, "zstd"): None}, + ((CodecDefinition, "zstd"),), + (), + ), + ( + CORE, + {(CodecDefinition, "zstd"): ZSTD_CODEC}, + (), + (Conflict((CodecDefinition, "zstd"), ZSTD_CODEC, None),), + ), + ( + Context.of(MY_GZIP), + {(CodecDefinition, "gzip"): GZIP_CODEC}, + (), + (Conflict((CodecDefinition, "gzip"), GZIP_CODEC, MY_GZIP),), + ), + (CORE, {(CodecDefinition, "acme.x"): None}, (), ()), + (CORE, {(DataTypeDefinition, "r*"): RAW_BYTES_DATA_TYPE}, (), ()), + ], + ids=["agrees", "gain", "loss", "conflict", "unclaimed-both", "raw-bits"], +) +def test_disagreements_says_where_a_scope_reads_claims_otherwise( + scope: Context, + claims: dict[Any, Any], + gains: tuple[Any, ...], + conflicts: tuple[Conflict, ...], +) -> None: + """A scope agrees with claims it reads identically, gains where it claims what the claims left unclaimed, and conflicts where it reads a name by another definition or by none.""" + found = scope.disagreements(claims) + assert isinstance(found, Disagreements) + assert (found.gains, found.conflicts) == (gains, conflicts) + assert found.agrees is (len(gains) == 0 and len(conflicts) == 0) + + +@pytest.mark.parametrize( + ("scopes", "joined"), + [ + ((CORE, CORE_AND_EXTENSIONS), CORE_AND_EXTENSIONS), + ((Context.of(GZIP_CODEC), Context.of(ZSTD_CODEC)), Context.of(GZIP_CODEC, ZSTD_CODEC)), + ((), Context.of()), + ((CORE, CORE), CORE), + ( + (Context.of(BYTES_CODEC), Context.of(BYTES_DATA_TYPE)), + Context.of(BYTES_CODEC, BYTES_DATA_TYPE), + ), + ], + ids=["subset", "disjoint", "none", "same", "same-name-two-kinds"], +) +def test_joined_is_the_least_scope_above_each(scopes: tuple[Context, ...], joined: Context) -> None: + """The join of scopes files every definition any of them files, once; one kind's name is not another's.""" + assert Context.joined(*scopes) == joined + + +def test_error_joined_refuses_one_name_filed_two_ways() -> None: + """Scopes that file different definitions under one name have no join: a `ScopeConflictError` naming the key and both definitions.""" + with pytest.raises(ScopeConflictError) as raised: + Context.joined(CORE, Context.of(MY_GZIP)) + (conflict,) = raised.value.conflicts + assert (conflict.key, conflict.claimed, conflict.found) == ( + (CodecDefinition, "gzip"), + GZIP_CODEC, + MY_GZIP, + ) + + +def test_claims_is_a_type_a_signature_can_hold() -> None: + """`Claims` resolves as an annotation at run time, as a caller's `get_type_hints` reads one: it is a type, not a string.""" + + def read(claims: Claims) -> None: + pass + + assert "claims" in get_type_hints(read) + + +@pytest.mark.parametrize( + ("kind", "said"), + [ + (CodecDefinition, "codec"), + (DataTypeDefinition, "data type"), + (ChunkGridDefinition, "chunk grid"), + (ChunkKeyEncodingDefinition, "chunk key encoding"), + (StorageTransformerDefinition, "storage transformer"), + ], + ids=["codec", "data-type", "chunk-grid", "chunk-key-encoding", "storage-transformer"], +) +def test_a_kind_is_named_in_words(kind: type[Definition[Any]], said: str) -> None: + """`kind_name` names each kind of definition as a message does: `ChunkKeyEncodingDefinition` is "chunk key encoding".""" + assert kind_name(kind) == said + assert said in str(Conflict((kind, "x"), None, None)) + + +def test_a_gain_is_judged_by_what_the_definition_reads_not_by_spelling() -> None: + """An accepted field refines an unclaimed one when its definition, reading what the unclaimed field wrote, reads the accepted field: two spellings of one configuration are one gain, so `refines` is transitive through `==`, and a spelling the definition reads otherwise is no gain.""" + from zarr_metadata.v3.definition import CORE, ChunkKeyEncodingDefinition, Context, resolve + + nothing = Context.of() + spelled_out, _ = resolve( + {"name": "default", "configuration": {"separator": "/"}}, ChunkKeyEncodingDefinition, CORE + ) + bare, _ = resolve({"name": "default"}, ChunkKeyEncodingDefinition, CORE) + unclaimed, _ = resolve({"name": "default"}, ChunkKeyEncodingDefinition, nothing) + assert spelled_out == bare + assert refines(bare, unclaimed) + assert refines(spelled_out, unclaimed) + other, _ = resolve( + {"name": "default", "configuration": {"separator": "."}}, ChunkKeyEncodingDefinition, CORE + ) + assert not refines(other, unclaimed) + shard = { + "name": "sharding_indexed", + "configuration": { + "chunk_shape": [2], + "codecs": [{"name": "bytes", "configuration": {"endian": "little"}}], + "index_codecs": [{"name": "bytes", "configuration": {"endian": "little"}}, "crc32c"], + "index_location": "end", + }, + } + read, _ = resolve(shard, CodecDefinition, CORE) + unread, _ = resolve({**shard, "must_understand": True}, CodecDefinition, nothing) + assert refines(read, unread) + + +OTHER_GZIP = CodecDefinition(name="gzip", configuration=Empty, kind="bytes_bytes", size="static") +"""A second definition under the core gzip's name: what a join conflicts on.""" + + +def test_a_scope_conflict_error_pickles_and_copies_with_its_conflicts() -> None: + """`ScopeConflictError` pickles and copies as it was raised: the same conflicts, the same message, as `MetadataValidationError` does, so a conflict reported in another process reads the same here.""" + import copy + + with pytest.raises(ScopeConflictError) as raised: + Context.joined(CORE, Context.of(OTHER_GZIP)) + error = raised.value + error.add_note("seen in a join") + for again in (pickle.loads(pickle.dumps(error)), copy.copy(error), copy.deepcopy(error)): + assert again.conflicts == error.conflicts + assert str(again) == str(error) + assert str(again).startswith("codec 'gzip'") + assert again.__notes__ == ["seen in a join"] + + +def test_a_conflict_names_the_field_as_the_document_writes_it_and_tells_the_definitions_apart() -> ( + None +): + """A conflict found in a document says the name the document writes, `r16`, not the name its definition is filed under, `r*`, and tells two definitions of one name apart by the configuration each declares, as the consolidated entries' messages do.""" + import dataclasses + + from zarr_metadata.model import ZarrV3ArrayMetadata + from zarr_metadata.v3.definition import CORE_AND_EXTENSIONS + + document = dict( + ZarrV3ArrayMetadata.create_default(shape=(2,), data_type="r16", fill_value=[0, 0]).to_json() + ) + model = ZarrV3ArrayMetadata(document, CORE_AND_EXTENSIONS) + other = dataclasses.replace(RAW_BYTES_DATA_TYPE, configuration=Empty) + with pytest.raises(ScopeConflictError) as raised: + model.refined_in(CORE_AND_EXTENSIONS.extended_with(other)) + message = str(raised.value) + assert message.startswith("data type 'r16'") + assert "RawBytesConfiguration" in message + assert "Empty" in message + with pytest.raises(ScopeConflictError) as joined: + Context.joined(Context.of(GZIP_CODEC), Context.of(OTHER_GZIP)) + assert str(joined.value).count("gzip") >= 2 + assert "Empty" in str(joined.value) + + +def test_error_a_scope_built_from_tables_keeps_the_invariants_of() -> None: + """`Context(tables)`, the constructor, refuses what `Context.of` refuses: a key that is no kind, and kinds of two formats, so no scope is built that `of` could not build.""" + from types import MappingProxyType + + from zarr_metadata.v2.data_type.scalar import UINT_V2 + from zarr_metadata.v2.definition import ZarrV2DataTypeDefinition + + with pytest.raises(TypeError, match="one Zarr format"): + Context( + MappingProxyType( + {CodecDefinition: {"gzip": GZIP_CODEC}, ZarrV2DataTypeDefinition: {"uint": UINT_V2}} + ) + ) + with pytest.raises(TypeError, match="kind"): + Context(MappingProxyType({int: {"gzip": GZIP_CODEC}})) # pyright: ignore[reportArgumentType] diff --git a/packages/zarr-metadata/tests/v3/test_sharding.py b/packages/zarr-metadata/tests/v3/test_sharding.py index 3afa98994b..bcf834b804 100644 --- a/packages/zarr-metadata/tests/v3/test_sharding.py +++ b/packages/zarr-metadata/tests/v3/test_sharding.py @@ -11,15 +11,20 @@ import pytest -from zarr_metadata.model import ZarrV3ArrayMetadata, validate_array_metadata_v3 +from zarr_metadata.model import ( + ZarrV3ArrayMetadata, + validate_array_metadata_v3, +) +from zarr_metadata.v3._pipeline import ( + read_pipeline, +) from zarr_metadata.v3.definition import ( CORE_AND_EXTENSIONS, Chunk, CodecDefinition, DataTypeDefinition, JSONValue, - Resolved, - read_pipeline, + ResolvedField, resolve, ) @@ -32,7 +37,7 @@ INDEX: list[JSONValue] = [LITTLE, "crc32c"] -def _dt(name: JSONValue) -> Resolved[DataTypeDefinition[Any]]: +def _dt(name: JSONValue) -> ResolvedField[DataTypeDefinition[Any]]: return resolve(name, DataTypeDefinition, CORE_AND_EXTENSIONS)[0] @@ -40,7 +45,9 @@ def _dt(name: JSONValue) -> Resolved[DataTypeDefinition[Any]]: UINT64 = _dt("uint64") -def _chunk(*axes: set[int] | None, data_type: Resolved[DataTypeDefinition[Any]] = FLOAT32) -> Chunk: +def _chunk( + *axes: set[int] | None, data_type: ResolvedField[DataTypeDefinition[Any]] = FLOAT32 +) -> Chunk: return Chunk(tuple(None if axis is None else frozenset(axis) for axis in axes), data_type) diff --git a/packages/zarr-metadata/uv.lock b/packages/zarr-metadata/uv.lock index 07ad9a0ebe..f1d80de5cc 100644 --- a/packages/zarr-metadata/uv.lock +++ b/packages/zarr-metadata/uv.lock @@ -1135,6 +1135,7 @@ wheels = [ name = "zarr-metadata" source = { editable = "." } dependencies = [ + { name = "annotated-types" }, { name = "typing-extensions" }, ] @@ -1155,7 +1156,10 @@ test = [ ] [package.metadata] -requires-dist = [{ name = "typing-extensions", specifier = ">=4.16" }] +requires-dist = [ + { name = "annotated-types", specifier = ">=0.6" }, + { name = "typing-extensions", specifier = ">=4.16" }, +] [package.metadata.requires-dev] docs = [