19 KiB
Borrow Time Machine's pressure-driven thinning discipline
Leading backup systems agree more than their docs admit: declarative time policy decides what may die, cheap metadata marking runs often, expensive reclamation runs rarely, capacity only forces the issue when the target fills, and local resource politeness comes from static ceilings plus OS deferral rather than true autotuning. For a local Linux file-versioning daemon syncing to WebDAV, that means copying Kopia's built-in maintenance rhythm, Time Machine's pre-backup thinning with padding, WebDAV PROPFIND quota checks, and a systemd slice with conservative defaults - because none of the surveyed tools offers a keep-usage-below-70% knob, and the one system that deletes on full does so without any percentage at all.
Cheap marking runs daily while reclamation waits weeks
The universal split is cheap expiration versus expensive reclamation, with very different cadences. In restic forget only deletes snapshot metadata while prune repacks data and is very time-consuming for remote repositories, bounded by --max-unused default 5% (Source). Borg made the same split explicit after 1.2 where repository disk space is not freed until you run borg compact and docs tell operators to run it regularly but not after every command, roughly monthly or when space is needed (Source). Borg's default compact --threshold 10% compacts only segments above that saving, with 2.x gating further to reclaimable space above threshold divided by five (Source). Kopia mirrors the design with quick maintenance hourly and full maintenance every 24h by default, where quick work never deletes metadata without another copy existing and full work does Snapshot GC plus pack compaction over several cycles (Source). Veeam applies short-term retention inline at the end of each job session while GFS deletions fall to a nightly background job, with health check verifying only the latest restore point per chain on a monthly default rather than whole history (Source).
Scheduling ownership falls on a spectrum from manual to fully automatic that the daemon should learn from. Restic and Borg ship no built-in scheduler at all, so retention becomes a wrapper convention of backup then forget plus prune or create then prune then compact, typically driven by daily cron such as 0 0 * * * with retention plus check monthly (Source). Borgmatic warns its every-run prune plus compact plus check default suits small repos but large repos must decouple them to separate schedules (Source. Kopia inverted this by making maintenance automatic since v0.6.0, firing opportunistically whenever the client runs with QuickCycle 1h and FullCycle 24h tunable via maintenance set and pausable via pause flags (Source). KopiaUI or server checks roughly every 10 minutes and does work only when due (Source. Time Machine goes furthest with zero user-visible retention scheduling, keeping hourly backups for 24 hours, daily for a month, weekly for previous months and thinning continuously plus under pressure (Source). Veeam sits in the enterprise middle with a per-job scheduler plus GFS calendar and monthly health-check piggybacked on the first session of the scheduled day (Source). Safety rails are strikingly convergent and worth copying wholesale: every tool offers dry-run previews without making them default, restic refuses empty policies and requires --unsafe-allow-remove-all plus filters to wipe a group, Borg retains the oldest archive when no rule keeps anything and since 2.x offers undelete until compact runs, and Kopia elects a single maintenance owner under exclusive lock with PackDeleteMinAge 24h so GC takes multiple cycles unless defeated with safety none (Source). Restic prune holds an exclusive lock that blocks backups while Borg and Kopia serialize through repo plus cache locks with break-lock and force overrides, and interrupted cheap passes rerun safely while expensive passes resume from checkpoints or recorded partial results (Source).
Time policy leads, capacity follows except when disk fills
Only Apple Time Machine makes capacity the primary retention driver, and even it leads with time. Backupd logs show the literal two-phase layering of Starting pre-backup thinning: 53.57 GB requested (including padding) then only if No expired backups exist - deleting oldest backups to make room (Source). The published ladder keeps first-of-day beyond 24 hours and first-of-week beyond 30 days, deleting oldest weeklies only when space is needed (Source). There is no published percentage watermark, deletion is on-demand pre-backup rather than a high-water mark (Source). Per-destination caps come only from tmutil setquota DESTINATION_ID QUOTA_IN_GB, which thins to fit on the next backup (Source). Local snapshots are purgeable space automatically reclaimed as needed with urgency-controlled manual relief via thinlocalsnapshots mount_point purge_amount urgency 1-4 (Source). Every other surveyed system treats capacity as monitoring, placement, or manual pruning rather than a retention input. Veeam retention defines the number of restore points to keep on performance and capacity extents, purging capacity-tier blocks on the next offload session, and when extents fill guidance is simply to add a new extent (Source). Veeam placement prefers the extent with fewest chains breaking ties by most free space, with priority always to complete a backup even by violating Data-Locality (Source). Forecast thresholds such as a worked Repository Free Space 30% example flag repositories that will run out rather than enforcing deletion (Source). S3 Lifecycle rules use Days, Date, and NoncurrentVersionExpiration per prefix or tag with no capacity trigger at all (Source). Hetzner Storage Box quotas are fixed plan sizes with 10, 20, 30, or 40 manual plus automatic snapshot slots consuming the same plan capacity alongside live data, offering no per-subaccount quota and no auto-thinning beyond setting directories read-only (Source). ZFS tooling is count-based with -k keep NUM recent snapshots and splits take, prune, and cron while capacity monitoring only feeds Nagios-style alerts (Source). The percentage numbers that do exist are health or performance floors, not auto-delete triggers, with Ubuntu warning that minimum free space to preserve ZFS performance is 20% while the pool sits at 10% (Source) and OpenZFS advising to keep pool free space above 10% to avoid metaslabs hitting 4% where the allocator collapses from first-fit to best-fit (Source). Borg likewise has no quota-aware prune and warns that repository disk space is not freed until you run borg compact, urging operators to ensure always plenty of free space and to prune plus compact regularly (Source).
| System | Primary retention driver | Capacity role | Published percentage trigger |
|---|---|---|---|
| Time Machine | Time ladder then oldest-first on full | Pre-backup thinning with padding | None, allocation-driven |
| Veeam SOBR | Restore-point counts plus GFS | Placement plus alerting, add extent | 30% forecast example only |
| S3 Lifecycle | Days and versions | None | None |
| ZFS, sanoid, Borg, restic | Counts and time windows | Monitoring plus manual prune | 20%, 10%, 4% health floors only |
Failure at 100% full explains why a headroom design matters more than any target percentage. Time Machine cancels the run with Stopping backup. Backup canceled. Compacting backup disk image to recover free space then retries as a fresh standard backup, surfacing This backup is too large for the backup disk. The backup requires XX GB but only YY GB are available when nothing is freeable (Source). Veeam guards production datastores with Skip VMs when free disk is below plus a hard 2 GB floor even when skipping is disabled (Source). Borg is the cautionary tale that if you do run out of disk space, it can be hard or impossible to free space, because Borg needs free space to operate - even to delete backup archives (Source). ZFS adds the snapshot-pinned trap where deleting a file present in a snapshot gains no space and can itself return unexpected ENOSPC (Source). Restic historically panicked on temp-pack writes with no space left on device and retried blindly on ENOSPC before rest-server fixes (Source). WebDAV at least fails cleanly with 507 Insufficient Storage on PUT, MKCOL, MOVE, or COPY (Source). The robust pattern combining all three headroom elements is pre-backup estimate plus padding as Time Machine does, a reserved-space tripwire as Veeam 2 GB, Borg repo-space, and ZFS 10% do, and a recovery path that works at 100% as Time Machine compact-and-retry does and Borg notably lacks. Crucially for snapshot-capable targets, a capacity pruner must delete snapshots themselves oldest-first rather than thinning live data, because Hetzner snapshots silently consume the same plan quota (Source).
WebDAV PROPFIND gives daemon a cheap capacity signal
WebDAV is unusually lucky among dumb backends because RFC 4331 standardizes quota discovery. Servers expose DAV:quota-available-bytes as maximum additional storage allocatable and DAV:quota-used-bytes including sub-resources, warning that as available bytes approaches zero further allocations may be refused (Source). Nextcloud implements both properties retrievable via PROPFIND (Source). That makes a Depth:0 PROPFIND for quota-used-bytes plus quota-available-bytes the cheapest correct pre-backup check for Hetzner or Nextcloud targets, avoiding any tree walk. Alternatives on other transports show why this matters. S3 has no byte-quota endpoint with accounting left to CloudWatch and Storage Lens and lifecycle evaluated by a daily asynchronous scan (Source). Hetzner exposes usage via console or API rather than uniformly in-protocol across its FTP, SFTP, rsync, Borg, SMB, and WebDAV paths (Source). Veeam polls extent free space but free space data is only retrieved when no active tasks are assigned to an extent, so placement under continuous load uses stale data (Source). Borg and restic on SSH, SFTP, or B2 do no server quota query at all, directing users to local filesystem monitoring and borg repo-space accounting (Source). Restic cache mismanagement shows the cost of local-only accounting when cache at system paths can grow to 20 to 40 GB or roughly 3 to 10% of repo size with no built-in cap (Source). Time Machine sidesteps the whole problem by asking the local filesystem how many bytes plus padding the next backup needs and thinning until that fits rather than tracking a remote percentage, a model the daemon should emulate by combining local size estimation with a PROPFIND confirmation before each sync.
Static ceilings plus OS deferral beat true autotuning
No surveyed tool probes link speed, disk size, or free RAM to size packs, connections, or limits. Rate knobs are consistently static, opt-in, and unlimited by default. Restic exposes --limit-upload and --limit-download in KiB per second defaulting to 0 meaning unlimited and cannot change them mid-run without restart (Source). Borg 1.x exposes --upload-ratelimit in kiByte per second default 0 over a token-bucket SleepingBandwidthLimiter with 0.1 second period capping bursts at twice the allowance (Source). Borg 2.x moved shaping to BORGSTORE_BANDWIDTH in bits per second default 0 plus BORGSTORE_LATENCY (Source). Kopia exposes repository throttle set with upload and download bytes per second plus concurrent reads, writes, and request rates (Source). Duplicity has no native generic limit with workarounds via trickle or tc (Source). Time Machine instead imposes mandatory kernel-level low-priority I/O where backupd runs at IOPOL_THROTTLE for long-running background work throttled to protect higher policies (Source). Disabling that global throttle lifted a measured copy phase from 193 to 332 MB per second and overall backup from 160 to 276 MB per second (Source). The only resource-derived defaults anywhere are CPU-count-derived worker pools. Restic defaults to file-read concurrency 2, blob-save concurrency equal to runtime.NumCPU, tree-save at 20 times blob concurrency, and backend connections 5 or 2 for local (Source). Kopia defaults --max-parallel-file-reads to logical core count and scales s2 compressor concurrency similarly (Source). Time Machine scheduling scores each due activity every few seconds against temperature and load via DAS-CTS, dispatching only above threshold, with Apple Silicon background threads confined to Efficiency cores at 972 to 1332 MHz and roughly 90% E-core residency capping throughput near 300 to 400 items per second (Source). Raising concurrency without memory awareness is dangerous, with one restic test at connections 16 and reads 16 hitting 300 GB RAM and OOM (Source). Constrained-box guidance therefore stays manual with cheap compression, lower parallelism, and larger packs. Borg defaults to lz4 with a proposal to move to zstd level -4 capped at 4 threads (Source. Kopia disables compression by default and benchmarks s2-default at 4 GiB per second and 375 MiB RSS versus zstd at 323 MiB per second on a 466 MiB corpus (Source). Restic on SMR USB disks converges on --read-concurrency 1 plus --pack-size 128 plus --no-cache with local connections 1, trading longer single-pack uploads and 64 to 384 MiB staging for fewer files (Source). Upstream tools ship no restrictive systemd units at all, leaving the enforceable layer to downstream slices. Systemd semantics give CPUQuota as percent of one CPU, CPUWeight 1 to 10000 default 100 for contention only, MemoryHigh as soft throttle versus MemoryMax as hard OOM kill, and IOWeight plus per-device bandwidth caps (Source). Community practice fills the gap with Nice 17 plus sandboxing and RandomizedDelaySec 300 on restic timers and CPUQuota 80% plus MemoryMax 2G on segmented Borg services (Source). Modern guidance recommends Nice 19 plus ionice class idle for soft priority with a backup.slice carrying CPUQuota 50 to 80%, MemoryHigh and Max 1 to 4G, IOWeight 10, and IOWriteBandwidthMax such as 30 MB per second, noting ionice is ignored on default mq-deadline SSD schedulers and MemoryMax kills rather than slows (Source).
Conclusion
The evidence reframes the daemon as a Kopia-style insider with Time Machine instincts rather than a restic-style script kit: own its scheduler with hourly cheap and daily full passes, enforce time policy first with capacity-triggered oldest-first eviction guarded by a minimum-retention floor, and treat PROPFIND plus local size-plus-padding as the single pre-sync gate instead of inventing a 70% watermark no vendor validates. Resource modesty then becomes configuration rather than cleverness, with static unlimited-by-default throttles, core-count workers, s2 or lz4 compression, and a systemd slice doing the throttling the backup code cannot do for itself.