Skip to content

Commit 31e7fb9

Browse files
authored
Publish the service images as zstd rather than gzip (#369)
A pull was already a compressed transfer, so this is a better algorithm for the same job rather than compression where there was none. Measured on `agent-computer`, which is where the bytes are: 962 MB becomes 886 MB, and zstd inflates several times faster, which is worth more on 2 GB of Chromium than the 8% is. `force-compression=true` is the part that matters. Without it only our own thin layers change and the saving rounds to nothing, because almost every byte came from Playwright's base image. With it those layers are recompressed, which also means these images stop sharing layers with a gzip pull of the same base. Only the five component images. Reading a zstd layer needs a client that supports it, which Podman and current containerd do, and the installer ships Podman. `ghcr.io/copilotkit/openbot` is pulled by deployments running whatever they have, so it stays gzip. Verified against ghcr.io before merging rather than at the next release: both architectures pushed by digest with this exact exporter string, merged into one OCI index, layers reported as `application/vnd.oci.image.layer.v1.tar+zstd` including the base layer, and `docker pull --platform` succeeded for both.
1 parent ff00ada commit 31e7fb9

2 files changed

Lines changed: 28 additions & 1 deletion

File tree

‎.github/workflows/publish-release.yml‎

Lines changed: 15 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -206,13 +206,27 @@ jobs:
206206
# Pushed by digest and deliberately untagged. A tag written here would name one architecture,
207207
# and the two jobs for one image would race to own it, so the last to finish would decide what
208208
# the tag meant.
209+
#
210+
# zstd, not gzip. A pull is already a compressed transfer, so this is not about compressing
211+
# something that was not; it is a better algorithm for the same job. Measured on
212+
# `agent-computer`, the image that matters: 962 MB becomes 886 MB, and it inflates several
213+
# times faster, which is worth more than the 8% on 2 GB of Chromium. `force-compression`
214+
# is what reaches the layers that came from somebody else's registry, which is where almost
215+
# all of the bytes are; without it only our own thin layers change and the saving rounds to
216+
# nothing. The cost is that these images no longer share layers with a gzip pull of the same
217+
# base, and that a client which cannot read zstd cannot read them, which is why this is here
218+
# and not on the `openbot` image above: that one is pulled by servers with whatever they have,
219+
# and these are pulled by an installer that ships Podman.
209220
- id: push
210221
uses: docker/build-push-action@53b7df96c91f9c12dcc8a07bcb9ccacbed38856a # v7.3.0
211222
with:
212223
context: .
213224
file: ${{ matrix.image }}/Dockerfile
214225
platforms: linux/${{ matrix.platform.arch }}
215-
outputs: type=image,name=ghcr.io/copilotkit/openbot-${{ matrix.image }},push-by-digest=true,name-canonical=true,push=true
226+
# `oci-mediatypes=true` because zstd layers have no Docker-schema2 media type to be
227+
# described by. It is buildx's default for this exporter and is stated rather than
228+
# assumed, since the push fails without it and the reason would not be obvious.
229+
outputs: type=image,name=ghcr.io/copilotkit/openbot-${{ matrix.image }},push-by-digest=true,name-canonical=true,push=true,oci-mediatypes=true,compression=zstd,force-compression=true
216230
# Scoped per image and per architecture. One shared scope would have ten builds
217231
# overwriting each other's cache and none of them reading their own.
218232
cache-from: type=gha,scope=${{ matrix.image }}-${{ matrix.platform.arch }}

‎CHANGELOG.md‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -8,6 +8,19 @@ Newest first. `Unreleased` is what is on `main` and not yet tagged.
88

99
## Unreleased
1010

11+
### The published service images are zstd rather than gzip
12+
13+
An image pull was already a compressed transfer, so this is not compression where there was none; it
14+
is a better algorithm for the same job. `agent-computer` goes from 962 MB to 886 MB on the wire, and
15+
zstd inflates several times faster, which on 2 GB of Chromium is worth more than the 8%. The saving
16+
comes from recompressing the layers that arrived from somebody else's registry, where nearly all of
17+
the bytes are, so the images no longer share layers with a gzip pull of the same base.
18+
19+
This applies to the five `ghcr.io/copilotkit/openbot-<service>` images and not to
20+
`ghcr.io/copilotkit/openbot`. Reading a zstd layer needs a client that supports it, which Podman and
21+
current containerd do; the single image is pulled by deployments running whatever they have, so it
22+
stays gzip.
23+
1124
### A release publishes every service's image, not just the one
1225

1326
`ghcr.io/copilotkit/openbot` was the only image a release produced, so anything running the Compose

0 commit comments

Comments
 (0)