20 KiB
Quota-Driven Retention
Which systems trigger retention from disk usage (e.g. "keep usage below X%", "delete oldest when over Y%"), and what exact thresholds do they use?
Takeaway
Only Apple Time Machine makes capacity the primary retention driver ("delete oldest when full"); Veeam, S3, Hetzner Storage Box, ZFS/sanoid and Borg/restic all use time/count-based retention as primary with capacity handled by monitoring, placement, or manual pruning - no "keep usage below X%" knob exists in most of them, and where percentage thresholds exist they are health-warning or performance floors (ZFS 20%/10%), not auto-delete triggers.
Cited Findings
- Time Machine "as your backup disk fills up, Time Machine deletes older backups to make room for new ones" with no published percentage - deletion is on-demand pre-backup, not a watermark - Source
- Time Machine backupd log shows two-phase scheme: "Starting pre-backup thinning: 53.57 GB requested (including padding)" then "No expired backups exist - deleting oldest backups to make room"; post-backup thinning separately expires hourly/daily backups - Source
- Time Machine standard thinning ladder is hourly backups >24h old (keep first-of-day), daily backups >30 days old (keep first-of-week), then oldest-weekly deletion only when space is needed - Source
- Time Machine per-destination quota via
tmutil setquota DESTINATION_ID QUOTA_IN_GBcaps a destination (e.g. 500 GB); quota "takes effect on the next backup, at which point older snapshots get thinned to fit" - Source tmutil thinlocalsnapshots mount_point [purge_amount] [urgency]: "tmutil will attempt (with urgency level 1-4) to reclaim purge_amount in bytes by thinning snapshots"; urgency 4 is most aggressive, e.g.thinlocalsnapshots / 10000000000 4to reclaim ~10 GB - Source; same signature confirmed in Apple developer thread - Source- Local APFS snapshots are purgeable space "automatically reclaimed as needed" but with no user-visible percentage; when reclamation lags users must thin manually (e.g.
thinlocalsnapshots / 100000000000 4≈ 100 GB at urgency 4) - Source - Veeam retention is count-based restore points / GFS, not capacity-based: "Retention policy defines the number of restore points to keep on your performance extents and capacity extents"; earliest restore point removed from chain, blocks purged from capacity tier on next offload/copy session - Source
- Veeam Scale-Out Backup Repository (SOBR) has no auto-delete-on-full: "if the extents of your scale-out backup repository run out of space, you can add a new extent"; free space on the new extent is added to SOBR capacity - Source
- Veeam SOBR placement prefers the extent with fewest chains, breaking ties by most free space; "priority is always to complete a backup" even if that violates the Data-Locality placement policy by spilling an incremental to another extent - Source
- Veeam ONE / MP capacity reports use a configurable "Repository Free Space (%)" forecast threshold (worked example 30%) to flag repositories that "will run out of space", i.e. monitoring/alerting rather than enforcement - Source
- S3 Lifecycle has no capacity trigger at all: rules are
Days/Date/NoncurrentVersionExpirationper prefix/tag (e.g. transition after 365 days, expire after 3650 days); S3 "quotas" are counts (buckets, access points), not bytes - Source; expiration example - Source - Hetzner Storage Box quota is the fixed plan size (BX11/BX21/BX31/BX41); snapshots "consume storage space from your Storage Box's storage capacity" alongside live data (
/.zfs/snapshot/), with slot caps of 10/20/30/40 manual + 10/20/30/40 automatic snapshots per plan - Source; plan overview confirms "unlimited traffic" but fixed storage - Source - Hetzner offers no per-subaccount quota ("there is currently no way to set quotas for each subaccount") and no auto-thinning; mitigation is manual read-only flag on sub-account directories - Source; official docs: "all sub accounts use the storage space of your Storage Box. To control storage usage, you can manually set a sub-account's directory to read-only" - Source
- ZFS tooling (zfs-auto-snapshot, sanoid) is count-based (
-k/--keep NUM Keep NUM recent snapshots), not usage-based; no--keep-below-X%option exists - Source; sanoid splits--take-snapshots/--prune-snapshots/--cronwith Nagios-style--monitor-capacityreporting only - Source - ZFS percentage numbers that do exist are health/performance floors, not retention triggers: Ubuntu warns "Minimum free space to take a snapshot and preserve ZFS performance is 20%. Free space on pool rpool is 10%" - Source; OpenZFS tuning advises "Keep pool free space above 10% to avoid many metaslabs from reaching the 4% free space threshold" where allocator flips from first-fit to best-fit and IOPS collapses - Source
- Borg has no quota-aware prune:
borg prune/borg delete+borg compactare explicit/manual or script-scheduled; "repository disk space is not freed until you run borg compact" - Source; quickstart warns to "ensure that there is always plenty of free space" and to "usepruneandcompactregularly" - Source
Inferences
- Time Machine is the outlier reference design for quota-aware retention: time ladder first, then unconditional oldest-first deletion driven by the byte size (plus padding) of the incoming backup.
- Every other surveyed system treats capacity as an ops/monitoring concern (add extent, raise quota, manual prune) rather than a retention input, which is why "keep usage below X%" knobs are absent outside custom wrappers.
Gaps
- No vendor-published numeric watermark (e.g. "start deleting at 90%") was found for Time Machine, Veeam SOBR, or Hetzner - deletion/placement appears driven by allocation failure or free-space comparison, not a fixed percent.
- Could not confirm any Hetzner-side automatic snapshot rotation on full; docs describe slot caps but not capacity-triggered eviction.
How do they measure remote usage on dumb backends (quota APIs, PROPFIND quota properties, du-style walks) when the server exposes no quota endpoint?
Takeaway
Only WebDAV-based backends have a standard quota API (RFC 4331 PROPFIND properties + HTTP 507); S3/object storage, Borg-over-SSH, restic, and Hetzner Storage Box over SFTP/rsync/Borg have no byte-quota endpoint, so clients fall back to local df/repository accounting, provider console/API, or expensive tree walks.
Cited Findings
- RFC 4331 defines two live PROPFIND properties for quota:
DAV:quota-available-bytes("maximum amount of additional storage available to be allocated") andDAV:quota-used-bytes("amount of space used ... including usage derived from sub-resources"), explicitly warning "as the DAV:quota-available-bytes on a resource approaches 0, further allocations ... may be refused" - Source - Quota exhaustion on WebDAV is signaled by HTTP 507 (Insufficient Storage), which "SHOULD be used when a client request (e.g. a PUT, PROPFIND, MKCOL, MOVE, or COPY) fails because it would exceed their quota or physical storage limits" - Source
- Nextcloud (a common self-hosted WebDAV target) implements both properties (
quota-available-bytes,quota-used-bytes) retrievable via PROPFIND - Source - Hetzner Storage Box exposes usage via console/API and a "Determine available Storage Box disk space" doc path, not via a uniform in-protocol quota on all transports; supported accesses are FTP/FTPS, SFTP/SCP, SSH/rsync/BorgBackup, SMB/CIFS, WebDAV - Source; only the WebDAV path inherits RFC 4331 properties.
- S3 has no byte-quota endpoint: lifecycle/quota docs cover object counts and bucket limits; storage accounting is via CloudWatch/Storage Lens/billing, and lifecycle evaluation is a daily asynchronous scan ("S3 Lifecycle evaluates objects against tag-based filters daily ... queues the action for asynchronous processing") - Source
- Veeam measures SOBR extent free space by polling the extent, but "Free space data is only retrieved when no active tasks are assigned to an extent", so placement decisions can be made on stale data under continuous load - Source
- Borg/restic on "dumb" backends (SSH, SFTP, rest-server, B2) do no server-side quota query; Borg docs direct users to local filesystem monitoring ("include the free space information in your backup log files"), client-side quotas, and
borg repo-spaceaccounting - Source - Restic cache-size issue threads show the failure of local-accounting fallback: cache at
~/Library/Caches/resticor/root/.cache/resticcan fill the system disk (reports of 20–40 GB, ~3–10% of repo size) with no built-in cap, andrestic cache --cleanuponly removes stale-repo caches - Source
Inferences
- For a WebDAV "dumb backend" (Hetzner, Nextcloud), PROPFIND
quota-used-bytes/quota-available-byteswith Depth:0 is the cheapest correct pre-backup check; on SFTP/rsync/S3 transports the only portable options are provider-specific APIs or recursive size walks (du,rclone size,restic stats), which are O(files) and unsuitable per-backup. - Time Machine's approach sidesteps remote accounting entirely: it asks the local filesystem how many bytes (plus padding) the next backup needs and thins until that fits, rather than tracking a remote percentage.
Gaps
- No evidence found that Borg, restic, or Veeam natively issue WebDAV PROPFIND quota checks; quota-awareness would have to be added in wrapper/scheduler code.
- Hetzner Storage Box per-protocol quota visibility (whether SFTP/SSH exposes the same numbers as WebDAV PROPFIND) is undocumented in fetched sources.
What is the recommended layering: time-based policy first, capacity trigger as backstop, or capacity as the primary driver?
Takeaway
Universal recommended layering is time/count policy first, capacity as backstop - except Time Machine, where capacity is the ultimate driver after the time ladder is exhausted. Enterprise guidance (Veeam, ZFS/sanoid, S3) never recommends capacity as the primary retention rule because it makes recovery windows unpredictable.
Cited Findings
- Time Machine layering: keep "local snapshots for the past 24 hours, daily backups for the past month and weekly backups for all previous months" and "oldest backups and any local snapshots are deleted as space is needed" - time ladder first, space-need second - Source
- Backupd implements the layering literally: pre-backup thinning first deletes expired (time-policy) backups, and only if "No expired backups exist" does it delete oldest backups to make room - Source
- Veeam layering: short-term + GFS (weekly/monthly/yearly) retention counts define what may be offloaded ("operational restore window ... defines which retention files can be offloaded"); capacity tier move/copy is placement, not an extra deletion rule, and "retention of the objects in the Capacity Tier is controlled by the backup or backup copy job's retention policy in restore points and not on repository level" - Source
- Veeam ONE guidance on low free space is "free up storage space on the repository or revise your backup retention policy" - i.e. human revises the time policy, system does not auto-shorten it - Source
- S3 layering is time-only by design: combine transition + expiration actions into a lifecycle timeline (e.g. 30d frequent → 90d infrequent → Glacier → expire); there is no capacity input to the rule engine - Source
- ZFS/sanoid layering is template keep-counts (hourly/daily/weekly/monthly) executed by
--cron, with--monitor-capacity/--monitor-healthfeeding external alerting, not feeding back into keep-counts - Source - Borg layering per quickstart: time/count
prunerules run on schedule pluscompactto actually reclaim, with free-space monitoring and optional reserved-space (borg repo-space) as the capacity backstop - Source
Inferences
- The sane default for a versioning/backup scheduler is: (1) declarative time-based keep rule, (2) pre-backup capacity check that prunes oldest-expired/oldest beyond-minimum first, (3) hard failure with clear "target full" signal rather than violating a minimum-retention floor silently.
- Time Machine's "delete oldest even if within time window" is acceptable for a single-user mirror but violates enterprise expectations (Veeam/S3) where immutability and GFS floors must survive capacity pressure.
Gaps
- No authoritative source found prescribing a specific headroom percentage (e.g. "always keep 10% free") as part of a retention policy; headroom numbers come from filesystem-tuning docs, not backup-policy docs.
Failure modes: what happens when the target is 100% full mid-backup, and how do systems reserve headroom?
Takeaway
Mid-backup-full behavior ranges from graceful (Time Machine aborts the run, compacts, retries; Veeam spills to another extent or skips VM below a free-space floor) to catastrophic (Borg may be unable to prune/compact without free space; ZFS deletions can themselves return ENOSPC when snapshots pin blocks; restic retries blindly on 507/ENOSPC in older versions).
Cited Findings
- Time Machine on unfreeable full: cancels the run ("Stopping backup. Backup canceled. Ejected Time Machine disk image. Compacting backup disk image to recover free space"), then retries as a fresh "Starting standard backup"; user-visible error is "This backup is too large for the backup disk. The backup requires XX GB but only YY GB are available" - Source; error text - Source
- Veeam datastore guard: jobs warn "Production datastore ... is getting low on free space (X GB left), and may run out of free disk space completely due to open snapshots" and "Skip VMs when free disk is below" logic terminates processing below the floor; hard floor is 2 GB free (registry
BlockSnapshotThreshold, DWORD GB) even if the skip option is disabled - Source - Veeam SOBR spillover: if one extent has no free space, Veeam places the next incremental on a different extent, violating Data-Locality to prioritize completing the backup - Source
- Borg worst case: "If you do run out of disk space, it can be hard or impossible to free space, because Borg needs free space to operate - even to delete backup archives"; mitigations are
borg repo-spacereservation, resizable LVs with unallocated extents, quotas, regular prune+compact - Source - Borg two-phase free: deleting an archive only marks for deletion; "repository disk space is not freed until you run borg compact" (which itself needs working space) - Source
- ZFS snapshot-pinned full: "if the file to be removed exists in a snapshot ... then no space is gained ... As a result, the file deletion can consume more disk space ... you can get an unexpected ENOSPC or EDQUOT when attempting to remove a file" - Source
- Restic mid-backup-full: local temp-pack path can panic with "no space left on device" (
panic: Write: write /tmp/restic-temp-pack-...: no space left on device) - Source; newer fix "Stop retrying uploads when rest-server runs out of space" shows prior behavior was unbounded retry on ENOSPC - Source - WebDAV full is a clean protocol error: 507 Insufficient Storage on PUT/MKCOL/MOVE/COPY - Source; Hetzner snapshots compound this because snapshot-pinned blocks silently consume the same plan quota - Source
- Headroom mechanisms found: Time Machine "padding" added to requested bytes in pre-backup thinning ("53.57 GB requested (including padding)") - Source; Veeam 2 GB snapshot floor - Source; Borg
repo-spacereservation + LVM overprovisioning - Source; ZFS 10–20% free-space guidance - Source
Inferences
- Robust design needs three headroom elements together: (a) pre-backup estimate + padding (Time Machine model), (b) a reserved-space tripwire that stops new writes before 100% (Veeam 2 GB / Borg repo-space / ZFS 10% models), (c) a recovery path that works at 100% (Time Machine compact-and-retry; Borg notably lacks one).
- Snapshot-pinned-full (ZFS, Hetzner Storage Box, APFS) is the nastiest failure mode: deleting live files does not free bytes, so a capacity-triggered pruner must delete snapshots themselves, oldest-first, not just thin live data.
Gaps
- Exact Time Machine padding formula and Veeam SOBR stale-free-space window under load are not published; both would need empirical measurement.
- No restic-side quota reservation feature found as of 2026 (open
--cache-size-limitrequest) - Source.