Skip to content

feat: stage and commit object writes with content ETags - #22

Open
sandros94 wants to merge 5 commits into
pithings:mainfrom
sandros94:fix/staged-commit
Open

sandros94 wants to merge 5 commits into
pithings:mainfrom
sandros94:fix/staged-commit

Conversation

@sandros94

Copy link
Copy Markdown

Closes #20, and the two things the thread found alongside it: a stale If-Match passing, and an unconditional PUT landing between a compare and its swap. Ofc I asked help from claude, especially with reviewing some ideas, benchmarking, docs and repo patterns, but everything is under my intent and knowledge. So feel free to mix, rewrite, ask or deny anything as usual @pi0

What changed

All three were the same write path, so there is now one: src/commit.ts, shared by mountx/s3 and mountx/webdav.

  1. A write is staged, then committed. The body streams to .mountx-multipart/tmp-<hex>, is checked there (length cap, Content-MD5), and becomes the object with link() for a create or an atomic rename() for a replace. A reader mid-PUT sees the whole previous object; a failed upload leaves it untouched and no debris. Capability-gated: memory and node-fs get this shape, unstorage writes in place as before
  2. The per-key lock covers the whole write, conditional or not, and deletes, markers, copy and multipart-complete destinations sit on the same chain. The roadmap entry accepting the unconditional-PUT race is gone
  3. The ETag is the content MD5, computed while the body streams and kept in a bounded per-bucket table validated by stat identity. Anything the table doesn't know (restart, eviction, an outside writer) answers the old derived -1 tag, so the worst case is one spurious 412, never a false 200. Multipart objects answer S3's md5(part md5s)-N; parts answer their MD5; Content-MD5 is verified on PutObject/UploadPart; CompleteMultipartUpload takes the same conditionals as PutObject
  4. Multigrain timestamps in the memory driver and unstorage's in-memory attributes, so their stat identity can't repeat inside one millisecond either

rclone check actually verifies hashes now instead of reporting them as unchecked.

Three things worth saying explicitly

.mountx-multipart should be renamed. As it now covers something broader than simple multi-part uploads, even just a simple .mountx-lock is clear enoguh IMO, but decision is yours

unstorage stays in place, and that is unstorage's limitation, not mountx's. It has no link and its rename is a copy followed by a delete, so there is nothing to commit atomically with. The stale-If-Match fix and the whole-write lock hold there too; what stays is that a reader can see a partial resource mid-write and a digest mismatch cannot restore the previous one. It's written down as such in the docs and in src/commit.ts.

The clock for the multigrain stamp is Date.now() plus a per-node nudge, discarding an earlier approach I did with performance.now(). The goal was a trusted wall clock that is runtime-agnostic, on older runtimes too: performance.now() reads a monotonic clock that excludes suspended time and drifts from the wall clock after a suspend or an NTP step. The nudge (previous stamp + 1 µs when the clock hasn't moved past it) encodes ordering, intentionally not elapsed time; nothing consumes the distance between two stamps as of writing this PR, and the module doc says so.

Measured

Stale If-Match accepted, through the session, same-size bodies, 200 rounds:

before after
node-fs, ext4, Linux 5.15 (Ubuntu 22.04) 177/200 0/200
node-fs, tmpfs, Linux 5.15 177/200 0/200
node-fs, vfat, Linux 6.18 200/200 0/200
memory / unstorage 183/200 / 182/200 0/200

Cost of the hash, 256 MiB PUT through the session on my dev machine: node-fs 747 → 667 MiB/s (S3) and 748 → 683 (WebDAV), the MD5 overlaps the threadpool write; memory 1205 → 590 on both, its write() is synchronous so copy and hash add up. A body paced like a socket (100 or 400 MiB/s) shows no loss on either driver. HEAD is unchanged on node-fs and faster on memory. Not on the docs pages but claude suggested correctly to place them in .agents/architecture.md with their provenance.

Verified

  • pnpm test green, plus the rclone oracle for both transports with rclone on PATH
  • the celld 0.5.0 fleet suite (test/integrations/celld.test.ts) on node-fs
  • the 5.15 numbers above come from a KVM guest with the official Ubuntu 22.04.5 cloud image

Commits

Five, meant to be read in order: the driver timestamps, the commit module with the S3 adoption, the WebDAV adoption, a small fix for 201/204 under a race (decided at the commit rather than before the lock), and the prose.

Follow-ups, deliberately out

  • x-amz-checksum-* headers and trailers (crc32/crc32c/sha1/sha256), noted in the roadmap; the plumbing is there
  • a mountx.* compare-and-commit extension for drivers whose store has a native CAS — the path to cross-process CAS, which this PR keeps out of scope (one gateway process per bucket)
  • the celld integration page still describes the old write path and links to an anchor this PR removes; I'll fix it once the pending revision of that page lands, to avoid two edits to the same file
  • R2 metadata compatibility, just an idea I want to validate and would fix one last gap to use mountx as celld's own S3.

@vercel

vercel Bot commented Sep 17, 2026

Copy link
Copy Markdown

@sandros94 is attempting to deploy a commit to the unjs Team on Vercel.

A member of the Team first needs to authorize it.

@pi0

pi0 commented Sep 17, 2026

Copy link
Copy Markdown
Member

Thanks for PR @sandros94 since it is a big change, can you somehow share your full claude chat session with me?

@sandros94

Copy link
Copy Markdown
Author

Thanks for PR @sandros94 since it is a big change, can you somehow share your full claude chat session with me?

Yes, absolutely. How would you like I send it to you? Discord?

@sandros94

Copy link
Copy Markdown
Author

I'm a bit second guessing myself on this topic, as ETags aren't durable and on process being restarted they fallback to the derivedETag. Which brings back the metadata topic again 🥲

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

s3: a PUT is visible half-written, so a reader mid-PUT sees a partial object

2 participants