Repository navigation
Store schemas as a compressed descriptor archive - #87
Merged
Merged
Conversation
psobot
force-pushed
the
psobot/historical-versions
branch
from
August 9, 2026 03:06
be464f3 to
85e8a57
Compare
Bundling eight Keynote versions took the wheel to 1.45 MB, 93% of it generated
_pb2.py. Two things inflate those: protoc embeds each schema's serialized
FileDescriptorProto as a Python bytes literal, where every non-printable byte
becomes a four-character \xNN escape - 6.48 MB of source around 3.10 MB of data
- and only 119 of the 258 version-file pairs are distinct, because Apple often
changes nothing in a given schema between releases.
Ask protoc for the FileDescriptorSet instead and store the descriptors
themselves: deduplicated, concatenated in filename order so near-identical
schemas compress against each other, and LZMA'd. Ordering is worth more than it
sounds - pickling the blobs separately gave 0.41 MB, concatenating them 0.09 MB.
wheel 1.45 MB -> 0.15 MB
archive 0.13 MB for all eight versions
which is smaller than the 0.21 MB single-version wheel this all started from,
and makes each further version nearly free.
Message classes come from message_factory against a per-version DescriptorPool,
so the pools that #81 introduced now fall out of the design rather than needing
generated code rewritten to use them. dumper/rewrite_imports.py and
dumper/rewrite_descriptor_pool.py are no longer needed.
Loading is also slightly faster: 12 ms to decompress the archive once, plus
~5 ms per version, against 33 ms to import one version's generated modules.
Two things found while building this:
- pool_for() held the module lock while calling _read_archive(), which takes
it too. A plain Lock deadlocked on the first pool built - it hung rather
than failed. Now an RLock, asserted directly by a test, because a test that
hangs is worse than one that fails.
- The type registry lived in the generated mapping modules, which this
replaces, and it cannot be derived from .proto files - it is extracted from
Keynote itself. It now lives in protos/versions/*/registry.json, checked in
beside the schemas it describes.
Also stops shipping the top-level dumper package, which had been landing in
site-packages since well before this change.
Verified byte-identical output against the generated-module implementation for
ls, cat, and unpack/pack under 10.2, 12.1 and 14.5. 227 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179M4xvAKPGrgKpy4AeCsM7
"descriptors.bin" said nothing about what the file holds. It is now protobuf_schemas.kpda: the stem names the contents, and the extension - Keynote Parser Descriptor Archive - matches the magic bytes already at the head of the file. Deliberately not ".xz" or ".lzma": the file is a container with its own header followed by two independently compressed sections, so naming it after the compressor would mislead anyone who tried to decompress it directly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0179M4xvAKPGrgKpy4AeCsM7
The custom format was a magic number, a length-prefixed JSON header, and a
hand-deduplicated blob addressed by offset. It worked, but nothing could read
it without this library, and it turns out it wasn't even buying anything.
The archive is now a .tar.xz holding one FileDescriptorSet per version - the
format protoc --descriptor_set_out already emits - plus registries.json:
$ tar tf keynote_parser/versions/protobuf_schemas.tar.xz
10.2.desc
...
registries.json
$ tar xOf protobuf_schemas.tar.xz 14.5.desc | protoc --decode_raw
Hand-deduplication turned out to be unnecessary. xz's window spans the whole
payload, so the schemas that repeat between versions - most of them - collapse
on their own. The result is 0.12 MB against the custom container's 0.13 MB,
despite storing 4.27 MB of descriptors rather than a deduplicated 1.62 MB.
Loading costs 30 ms rather than 17 ms, because there is more to decompress.
That is once per process, and still below the 33 ms the generated modules took.
Removes the hashing, the span index, the magic-number framing and the manual
ordering pass. Tar metadata is fixed so the archive is reproducible: two builds
from the same inputs are byte-identical, verified.
Output remains identical to the generated-module implementation for ls, cat and
unpack/pack under 10.2, 12.1 and 14.5. 227 tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179M4xvAKPGrgKpy4AeCsM7
Replacing the generated mapping modules left dumper/run.py --app-path still writing the old kind: a mapping.py importing a .generated package that no longer exists. Adding support for a new Keynote version would have produced artifacts nothing could load - the one path that matters most for this project's future, and the one with no test because it needs a debugger and an installed Keynote. It now writes what the archive expects: the extracted registry to protos/versions/<version>/registry.json, beside the protos it describes, and a mapping.py that is the same shim every other version has. The schemas themselves go into the shared archive in step 5. dumper/generate_mapping.py has no callers left and is removed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0179M4xvAKPGrgKpy4AeCsM7
psobot
force-pushed
the
psobot/descriptor-archive
branch
from
August 9, 2026 03:08
5ee9512 to
9fa2428
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This reduces code duplication in the published wheel.