Skip to content

feat: Bytes partial decode - #4458

Open
bendichter wants to merge 6 commits into
zarr-developers:mainfrom
bendichter:bytes-partial-decode
Open

bendichter wants to merge 6 commits into
zarr-developers:mainfrom
bendichter:bytes-partial-decode

Conversation

@bendichter

Copy link
Copy Markdown
Contributor

Summary

I am working with an application where data is stored uncompressed, with BytesCodec as the only codec, and I want to read small subregions of large chunks. Currently, zarr-python reads the entire chunk for every selection. That is necessary when a chunk is compressed, but not when BytesCodec is the only codec. This PR makes BytesCodec implement partial decoding, so it fetches only the range that corresponds to the rows of interest. Selections that touch every row, and chunks with any other codec, are read whole as before.

This helps any uncompressed array with large chunks, and especially virtual datasets, where the chunk layout comes from existing files. A contiguous HDF5 dataset, for example, becomes a single chunk that can be many GB. In an example I am working on, reading 100 samples from a 20 MB single-chunk array fetched 20 MB before this change and 2 KB after.

Author attestation

  • I am a human, these are my changes, and I have reviewed and understood every change and can explain why each is correct.

TODO

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.md
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

bendichter and others added 5 commits September 30, 2026 01:25
BytesCodec now implements partial decoding. An uncompressed chunk is
stored in C order, so each row along its first axis is a contiguous run
of bytes; for a selection that does not touch every row, the rows from
the first to the last one it touches are fetched with a single range
request and the selection is applied to them. The pipeline already
uses partial decoding when the array-to-bytes codec supports it and
there are no array-to-array or bytes-to-bytes codecs, so this applies to
uncompressed arrays only. Reading 100 rows of a single-chunk 20 MB
array now fetches the bytes of those rows instead of the whole chunk.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… reads

A boolean mask on the first axis, as from arr.oindex[mask], was trimmed to
start at the first selected row but not to end at the last, so its length
no longer matched the rows fetched and numpy raised an IndexError. The row
window and the shifted selection are now computed together in _row_window,
which trims the mask at both ends. This also gives mypy the narrowing it
needs.

FsspecStore over HTTP returns the whole object when a server ignores the
Range header. The partial read then failed to reshape a chunk that the full
read would have decoded, so a read that worked before this branch raised an
error. A response longer than the requested range is now sliced to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added the needs release notes Automatically applied to PRs which haven't added release notes label Sep 30, 2026
@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.11111% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 94.46%. Comparing base (d18fb50) to head (42871ff).

Files with missing lines Patch % Lines
src/zarr/codecs/bytes.py 91.11% 4 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main    #4458      +/-   ##
==========================================
- Coverage   94.47%   94.46%   -0.02%     
==========================================
  Files          93       93              
  Lines       13241    13284      +43     
==========================================
+ Hits        12510    12549      +39     
- Misses        731      735       +4     
Files with missing lines Coverage Δ
src/zarr/codecs/bytes.py 96.00% <91.11%> (-2.79%) ⬇️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Comment thread src/zarr/codecs/bytes.py
Comment on lines +174 to +177
if len(chunk_bytes) > (stop - first) * row_bytes:
# The store sent the whole chunk, as an HTTP server that ignores
# the Range header does.
chunk_bytes = chunk_bytes[first * row_bytes : stop * row_bytes]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this needs a test. if the store honored the byte range start but not the stop, then re-indexing from the start again will generate an invalid result.

because byte range handling is so important, we should probably set up stores in our test fixtures that span the range of byte range handling behavior, to ensure that branches like this get tested thoroughly

@d-v-b

d-v-b commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

thank you for this! would you mind implementing the same behavior for a synchronous method (_decode_partial_sync)? This would allow the FusedCodecPipeline to use this feature.

@d-v-b d-v-b added the benchmark Code will be benchmarked in a CI job. label Oct 1, 2026
@codspeed

codspeed Bot commented Oct 1, 2026

Copy link
Copy Markdown

Merging this PR will improve performance by 22.82%

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 33 improved benchmarks
✅ 80 untouched benchmarks
⏩ 37 skipped benchmarks1

Performance Changes

Benchmark BASE HEAD Efficiency
⚡ test_sharded_morton_indexing_large[(30, 30, 30)-memory] 7.6 s 5.5 s +39.45%
⚡ test_sharded_morton_indexing[(32, 32, 32)-memory] 1,160.7 ms 832.8 ms +39.38%
⚡ test_sharded_morton_indexing_large[(32, 32, 32)-memory] 9.2 s 6.6 s +39.1%
⚡ test_sharded_morton_indexing_large[(33, 33, 33)-memory] 10.1 s 7.3 s +39.1%
⚡ test_slice_indexing[None-(slice(None, None, None), slice(None, None, None), slice(None, None, None))-memory] 355.8 ms 256.4 ms +38.74%
⚡ test_sharded_morton_indexing[(16, 16, 16)-memory] 145.7 ms 105.1 ms +38.59%
⚡ test_slice_indexing[None-(slice(None, None, None), slice(None, None, None), slice(None, None, None))-memory_get_latency] 411.3 ms 307.2 ms +33.9%
⚡ test_slice_indexing[(50, 50, 50)-(slice(None, None, None), slice(None, None, None), slice(None, None, None))-memory] 406.2 ms 307.5 ms +32.12%
⚡ test_slice_indexing[(50, 50, 50)-(slice(None, None, None), slice(None, None, None), slice(None, None, None))-memory_get_latency] 408.5 ms 310.4 ms +31.6%
⚡ test_slice_indexing[None-(slice(None, None, None), slice(0, 3, 2), slice(0, 10, None))-memory] 3.4 ms 2.6 ms +31.34%
⚡ test_slice_indexing[None-(slice(None, None, None), slice(0, 3, 2), slice(0, 10, None))-memory_get_latency] 4 ms 3.1 ms +30.5%
⚡ test_read_array[latency=0-batched-local-Layout(shape=(1000000000,), chunks=(100000,), shards=(10000000,))-None-repeated] 4.4 s 3.5 s +23.82%
⚡ test_read_array[latency=0-batched-local-Layout(shape=(1000000000,), chunks=(100000,), shards=(10000000,))-None-semi_random] 4.4 s 3.5 s +23.46%
⚡ test_read_array[latency=0.03-batched-memory-Layout(shape=(1000000000,), chunks=(100000,), shards=(10000000,))-None-repeated] 4.4 s 3.7 s +20.89%
⚡ test_read_array[latency=0.03-batched-memory-Layout(shape=(100000000,), chunks=(100000,), shards=(100000000,))-None-repeated] 508.1 ms 420.9 ms +20.72%
⚡ test_read_array[latency=0.03-batched-memory-Layout(shape=(1000000000,), chunks=(100000,), shards=(10000000,))-None-semi_random] 4.4 s 3.7 s +20.28%
⚡ test_read_array[latency=0.03-batched-memory-Layout(shape=(100000000,), chunks=(100000,), shards=(100000000,))-None-semi_random] 503 ms 419.6 ms +19.88%
⚡ test_slice_indexing[None-(slice(None, 10, None), slice(None, 10, None), slice(None, 10, None))-memory] 868.2 µs 725.7 µs +19.64%
⚡ test_read_array[latency=0-batched-local-Layout(shape=(100000000,), chunks=(100000,), shards=None)-None-repeated] 671.7 ms 565.8 ms +18.71%
⚡ test_read_array[latency=0-batched-local-Layout(shape=(100000000,), chunks=(100000,), shards=None)-None-semi_random] 668.1 ms 564.9 ms +18.27%
... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing bendichter:bytes-partial-decode (1dfc3f3) with main (47528fe)2

Open in CodSpeed

Footnotes

  1. 37 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on main (d8c07e1) during the generation of this report, so 47528fe was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmark Code will be benchmarked in a CI job. needs release notes Automatically applied to PRs which haven't added release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants