Repository navigation
[Feature Request] Display size comparison to PBS WebUI #19
Description
Activity
Thank you for your contribution to this project.
This is a very good question and a perfect example of where the PBS GUI and the chunk-based view show different things 🙂
The PBS_Chunk_Checker works purely on chunk digests and answers the question: How much physical datastore space is required for the selected scope?
The scope can be a single restore point, a whole VM, multiple VMs or even a complete namespace.In the chunk usage summary:
unique are chunks that are referenced exactly once inside the selected scope.
In your case you checked only one backup of a single host, therefore every chunk is used only once andunique == total size.
This value represents the real datastore space this backup would consume if it were stored alone on a PBS without any other backups.duplicate are chunks that occur multiple times inside the selected scope and are therefore deduplicated.
duplicate ref chunks = 0means there are no identical chunks inside this backup, so nothing inside this scope benefits from deduplication. This backup does not deduplicate against itself.total ref is simply the sum of unique and duplicate chunk references.
The size shown in the PBS GUI is a logical size, not the physical datastore usage.
For a VM this is basically the sum of all configured virtual disk sizes. A VM in PVE with two 500 GB disks will be shown as 1 TB in the GUI even if only 200 GB per disk are actually used (thin provisioning). PBS shows the same logical value for the backup.So in your example:
GUI size = 2.1 TiB → logical/provisioned VM disk size
PBS_Chunk_Checker unique ≈ 1.2 TiB → real stored data / physically required spaceThis does not mean that deduplication saved you ~1 TB.
It simply means the VM disks are not filled completely and only about 1.2 TiB of real data exist.Because you checked only a single backup and
duplicate = 0, there are no deduplication savings inside the checked scope.For a first cloud sync to an empty remote datastore you would need to transfer approximately the unique size (~1.2 TiB), because only unique chunks have to be uploaded. Further backups would then benefit from deduplication.
In short: the GUI shows the logical disk size, while the script shows the real physical datastore usage for the selected scope.
Thanks @VoltKraft for the detailed response!
On a side note, this was a backup taken with proxmox-backup-client from a physical disk, e.g. an internal SSD which is used as NAS, not really a VM. That disk is 4TB big and is at 2.1Tb usage with data.
So if I understand this correctly, the difference shown here (1.2TB vs 2.1TB, could come from deduplication, but in my case is not, because there are no duplicate chuncks in a single backup? This also seems a bit weird, as I know there are big (10GB+) files that are in different folder but exactly the same file blob, which should be duplicate "blocks" or did I just have no luck, and no block matched here?
The interesting thing is, this is a backup of a single SSD, which currently sits at 2.1 of 3.6TB Data usage, so I'm kinda wondering where the 1.2TB usage come from? Does this factor in the compression,
In short: the GUI shows the logical disk size, while the script shows the real physical datastore usage for the selected scope.
So then again, would it be possible to show the logical size with the script too to see comparison to the size reported by proxmox backup server?If I may ask (no blame) but was that an LLM-Generated response?
Thanks for the clarification regarding the host backup — I initially missed that detail in the screenshot.
To be honest, my personal experience so far is mostly with VM and container backups, not with full host filesystem backups viaproxmox-backup-client.Proxmox Backup Server applies several techniques that influence the stored backup size. Thin provisioning and compression of the data. I therefore assume that the reduction we see here (from ~2.1 TiB filesystem usage to ~1.2 TiB physically stored chunk data) mainly comes from:
- only backing up actually used data (not empty space)
- compression of the chunk contents
According to my understanding of PBS, chunk deduplication happens on top of that. In your current test you are looking at a single snapshot of a single system, so there is not much opportunity yet for cross-snapshot deduplication. That matches the reported
0 duplicate ref chunks.Once you start looking at multiple restore points of the same system, you should see a significantly higher deduplication rate, because unchanged data will reference already existing chunks in the datastore.
The difference you are seeing is expected for a first backup of a single system, and mathes some of my own tests.
For example, I just did a quick test with my Nextcloud CT. With 2.25 TiB of storage used, I got 1.4 TiB and 0 duplicates in the PBS datastore with only one restore point checked. However, when I checked the CT across a total of 182 restore points, I got 1.4 TiB and 44973416 duplicates (>98%). (There aren't many changes in my Nextcloud, which is why there are so many duplicates.)
Deduplication benefits will mainly become visible when comparing multiple backups over time.And regarding your last question — guilty as charged 🙂: I usually dictate these topics into my phone in a rather unstructured way, and LLMs are extremely helpful to turn that into a properly structured and readable response with a bit more context.
Okay I see yeah I guess that makes sense.
thanks for the exaplanations- Repository owner locked and limited conversation to collaborators
on Mar 8, 2026
Hi, Thanks for building this awesome script!
This is exactly what I searched for!
Is your feature request related to a problem? Please describe.
I wanted to see, how much disk space a big backup is actually using on my datastore, to roughly estimate, how long a cloud backup would need to run with my WAN speed.
Describe the solution you'd like
It would be really nice, if the script showed the size, that PBS reports in the Content Section for this Backup.
Describe alternatives you've considered
Additional context
Here an example for the same backup:
So I assume this means, that the Deduplication saved me 1 TB and if I do a cloud backup it would need to upload roughly 1.2TiB