Files
versioning/research_notes/Retention scheduling adaptation/resource-adaptation.md
T

103 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Resource Adaptation in Backup Tools
## Which knobs exist (restic --limit-upload, borg --upload-ratelimit, kopia throttling, Time Machine low-priority I/O) and what are their defaults?
### Takeaway
All Linux tools use static, opt-in, unlimited-by-default rate/priority knobs; Time Machine instead imposes mandatory kernel-level low-priority I/O (IOPOL_THROTTLE) with no per-backup user knob.
### Cited Findings
- restic exposes global `--limit-upload rate` and `--limit-download rate` in KiB/s, default unlimited (0) - [Source](https://restic.readthedocs.io/en/stable/manual_rest.html); also documented in man pages as `--limit-upload=0` / `--limit-download=0` default unlimited - [Source](https://man.archlinux.org/man/restic-options.1.en)
- restic documents that `--limit-upload` cannot be changed mid-run without restart; users request pv-like `-R` dynamic adjustment and resort to killing/restarting or external shapers - [Source](https://forum.restic.net/t/change-limit-upload-without-aborting-restic/5958)
- Borg 1.x exposes `--upload-ratelimit RATE` in kiByte/s, default 0=unlimited (with `--remote-ratelimit` deprecated alias) - [Source](https://manpages.debian.org/bookworm/borgbackup/borg-common.1.en.html); also `--upload-buffer` size in MiB, default no buffer - [Source](https://man.archlinux.org/man/borg-common.1.en.txt)
- Borg 2.x removed `--remote-ratelimit`/`--upload-ratelimit`; bandwidth is limited via `BORGSTORE_BANDWIDTH` in bits/sec, default 0=unlimited, plus `BORGSTORE_LATENCY` delay per call, or via `pv -L` ProxyCommand / rclone `--bwlimit` - [Source](https://borgbackup.readthedocs.io/en/latest/faq.html)
- Older Borg 1.x FAQ documents upload-only `--remote-ratelimit` plus `pv`-wrapper + `BORG_RSH` for download shaping with on-the-fly `pv -R $(pidof pv) -L` changes - [Source](https://borgbackup.readthedocs.io/en/stable/faq.html)
- Borg's limiter is a token-bucket `SleepingBandwidthLimiter` with `RATELIMIT_PERIOD = 0.1`, capping burst quota at 2x period allowance - [Source](https://github.com/borgbackup/borg/blob/da3105f1/src/borg/remote.py)
- Kopia exposes `repository throttle set` flags `--upload-bytes-per-second`, `--download-bytes-per-second`, `--concurrent-reads/writes`, `--read/write-requests-per-second`, `--list-requests-per-second` - [Source](https://kopia.io/docs/reference/command-line/common/repository-throttle-set/); readable via `repository throttle get` - [Source](https://kopia.io/docs/reference/command-line/common/repository-throttle-get/); server-side equivalent `server throttle set` - [Source](https://kopia.io/docs/reference/command-line/common/server-throttle-set/)
- Kopia direct-connect backends accept `--max-upload-speed`/`--max-download-speed` bytes/sec at `repository connect` time (e.g. B2/S3), persisted as `maxUploadSpeedBytesPerSecond` in repository.config - [Source](https://kopia.discourse.group/t/limit-upload-speed-as-a-policy/990)
- Kopia `repository throttle set` is rejected on server-connected repos ("operation supported only on direct repository"), and per-policy/UI global throttle was still missing as of 2024–2026 feature requests - [Source](https://github.com/kopia/kopia/issues/3051); KopiaUI throttle exposure requested - [Source](https://github.com/kopia/kopia/issues/3586)
- Kopia upload throttling was bursty (whole-object sleeps) until PR #2682 added per-read `DuringUpload` throttling - [Source](https://github.com/kopia/kopia/pull/2682)
- Time Machine's `backupd` runs at `IOPOL_THROTTLE`, defined as "long-running I/O intensive background work, such as backups" that "will be throttled to prevent impact on higher policy levels" - [Source](https://eclecticlight.co/2022/01/20/why-time-machine-backups-can-be-interminably-slow/)
- Full `IOPOL` ladder is IMPORTANT (default) / STANDARD / UTILITY / THROTTLE / PASSIVE - [Source](https://eclecticlight.co/2026/03/28/explainer-i-o-throttling/)
- Global kill-switch `sudo sysctl debug.lowpri_throttle_enabled=0` (re-enable with `=1`, lost on reboot unless persisted via `/etc/sysctl.conf` or LaunchDaemon) removes throttle for all background I/O, not just backupd - [Source](https://osxdaily.com/2016/04/17/speed-up-time-machine-by-removing-low-process-priority-throttling/); same command/LaunchDaemon recipe - [Source](https://apple.stackexchange.com/questions/181609/time-capsule-wired-backup-transfer-slow-with-fast-bursts); throttling is I/O not CPU - [Source](https://mjtsai.com/blog/2016/03/16/massively-speed-up-time-machine-backups/)
- Measured effect of disabling throttle: copying phase 193→332 MB/s, overall backup 160→276 MB/s (>10 GB test); pre-backup 50 MB probe writes unaffected - [Source](https://eclecticlight.co/2022/02/28/does-removing-i-o-throttling-make-backups-faster/)
- Duplicity has no native generic bandwidth-limit option (open bug #1291633); workarounds are `trickle -s -u/-d`, WonderShaper/tc, or legacy `--scp-command="scp -l N"` (kbit/s, scp backend only, option later deprecated) - [Source](https://bugs.launchpad.net/bugs/1291633); scp `-l` throttle Q&A - [Source](https://lists.libreplanet.org/archive/html/duplicity-talk/2007-09/msg00058.html); router-QoS/DSCP attempts reported ineffective, per-machine Bandwidth Limiter used instead - [Source](https://lists.libreplanet.org/archive/html/duplicity-talk/2021-09/msg00000.html)
- Neither restic, borg, kopia, nor duplicity ships battery- or metered-network-aware auto-pause; scheduling/power-awareness is delegated to systemd timers, DAS-CTS (macOS), or external shapers (see Gaps).
### Inferences
- Rate limiting is consistently an operator-set static ceiling, not adaptive congestion control; only pv/WonderShaper/BORGSTORE allow out-of-band changes without restarting the backup.
- Borg 2.x shifts shaping out of borg CLI into the borgstore layer, unifying ssh/sftp/rclone paths but breaking old `--upload-ratelimit` scripts.
### Gaps
- No reliable source found for a built-in metered-WiFi or on-battery auto-suspend in restic/borg/kopia/duplicity as of 2026; Time Machine power/thermal inputs are internal to DAS scoring, not user flags.
## Does any tool auto-tune from measured resources (disk size, RAM, link speed) rather than static config, and how?
### Takeaway
No tool auto-tunes from measured disk/RAM/link speed; the only resource-derived defaults are CPU-count-derived worker counts, everything else is fixed static defaults.
### Cited Findings
- restic defaults: file-read concurrency 2 ("sweet spot" from HDD experiments), blob-save concurrency = `runtime.NumCPU()`, tree-save concurrency = 20x blob concurrency - [Source](https://github.com/restic/restic/blob/de9136b29f86216bd3e41397d19b25f26b578833/internal/archiver/archiver.go)
- restic uses all available CPUs by default; `GOMAXPROCS=1` pins to one core and slightly reduces memory - [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html)
- restic backend connection limit defaults to 5 (2 for local backend), tunable via `-o rest.connections=5` / `-o local.connections=2`; too-high values degrade performance - [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html)
- restic `--read-concurrency` / `RESTIC_READ_CONCURRENCY` raises parallel file reads for NVMe; `--no-scan` skips the pre-backup file-count/size scan that costs extra I/O on network/FUSE mounts - [Source](https://restic.readthedocs.io/en/latest/047_tuning_parameters.html); read-concurrency flag added in 0.15.0 for fast storage - [Source](https://restic.net/blog/2023-01-12/restic-0.15.0-released/)
- Single-large-file chunking in restic is sequential per file (~450 MB/s) with parallel hash/compress/encrypt downstream; multi-file parallelism is what scales, which is why `cores/4`-style read-concurrency guesses only hold for SSDs - [Source](https://github.com/restic/restic/issues/4477)
- Kopia `--max-parallel-file-reads` defaults to number of logical CPU cores; lowering it lowers CPU at cost of time - [Source](https://kopia.io/docs/faqs/)
- Kopia `--max-parallel-snapshots` controls simultaneous snapshots (server/KopiaUI); s2 `default`/`better` compressor concurrency equals logical core count, `s2-parallel-4/8` pins it - [Source](https://kopia.io/docs/advanced/compression/)
- Time Machine scheduling via DAS-CTS scores each due background activity every few seconds against temperature, load, and priority, dispatching `com.apple.backupd-auto` via XPC only when score exceeds threshold - [Source](https://eclecticlight.co/2023/11/28/scheduling-and-dispatch-of-backups-and-other-background-activities/)
- On Apple Silicon, background-QoS `backupd` threads are confined to Efficiency cores at reduced frequency (~972–1332 MHz, ~90% residency on E cores), capping throughput at ~300–400 items/s regardless of queue depth - [Source](https://eclecticlight.co/2022/01/20/why-time-machine-backups-can-be-interminably-slow/)
- No evidence any tool probes link speed, disk size, or free RAM to set pack size, connections, or limits automatically; restic 16 MiB pack size, Kopia parallelism, and Borg rate defaults are all static.
### Inferences
- "Auto-tuning" in this space means CPU-count-proportional worker pools plus OS-scheduler deferral (DAS-CTS/QoS), not closed-loop adaptation to throughput or memory pressure.
- Raising concurrency without raising memory (restic 100×1 GB test hit 300 GB RAM at connections=16/reads=16 and OOMed) shows why static defaults stay conservative - [Source](https://github.com/restic/restic/issues/4477).
### Gaps
- No source found documenting link-speed probing or disk-size-derived chunk/pack sizing in any of the five tools; if it exists it is not in public docs/CLI help.
## How do they behave on constrained boxes (low RAM SQLite/index handling, single-core compression choices)?
### Takeaway
Constrained-box guidance is manual: pick cheap compression, lower parallelism/connections, enlarge packs, and accept slower runs; low-RAM index handling remains a known failure mode, not an auto-degraded mode.
### Cited Findings
- Borg compression default is lz4 (very high speed, very low compression); alternatives `zstd[,L]` (default level 3), `zlib`, `lzma`, `auto,`, `none` - [Source](https://borgbackup.readthedocs.io/en/stable/usage/help.html); usage examples recommend `zlib,6` for ratio at cost of speed - [Source](https://borgbackup.readthedocs.io/en/stable/usage/create.html)
- Upstream Borg PR proposes moving default from `lz4` to `zstd,-4` (multithreaded, as-fast-or-faster creates, slightly better ratio; large incompressible-image corpora stay ~11% slower) with MT workers capped at 4 - [Source](https://github.com/borgbackup/borg/pull/10100)
- Kopia compression is disabled by default and set per-policy via `kopia policy set [--global] --compression=<...>` with min/max-size gates - [Source](https://kopia.io/docs/faqs/); full option list includes `s2-default/better/parallel-4/8`, `zstd/zstd-fastest/better`, `gzip/pgzip/deflate` variants - [Source](https://kopia.io/docs/reference/command-line/common/policy-set/)
- Kopia FAQ names compression + parallelism as the two main memory culprits; recommends disabling compression or using `s2`/`deflate`/`gzip` on small files under low memory, and lowering `--max-parallel-snapshots` / `--max-parallel-file-reads` - [Source](https://kopia.io/docs/faqs/)
- Kopia benchmark table (466 MiB corpus): s2-default 4 GiB/s at ~375 MiB RSS vs zstd 323 MiB/s at ~238 MiB vs zstd-best 19 MiB/s; on tiny files s2 stays fastest with ~2 MiB footprint - [Source](https://kopia.io/docs/advanced/compression/)
- Kopia zstd levels map to upstream klauspost/compress: fastest≈1, default≈3, better≈7, best≈11; higher custom levels (e.g. -22 --ultra --long) require code change - [Source](https://kopia.discourse.group/t/what-is-zstd-best-compression-and-can-i-customize/4934)
- restic on constrained HDD/NVMe: forum-tested recipe for 2.5 TB SMR-USB run is `--read-concurrency 1 --pack-size 128 --no-cache` (pack default 16 MiB raised to reduce file count and per-pack latency), with `local.connections=1` suggested for vibration-sensitive HDDs - [Source](https://forum.restic.net/t/first-backup-2-5tb-50-hours-can-i-improve-it/8288)
- restic large-pack tradeoff: bigger packs reduce file count and help HDD/Swift/Drive limits but need more `$TMPDIR` staging (64–384 MiB guidance) and longer single-pack uploads that wear SSDs - [Source](https://github.com/restic/restic/blob/master/doc/047_tuning_parameters.rst)
- Borg on near-full disks: 1.5 GiB-free VM case shows Borg needs headroom for segments/cache/index; workarounds discussed are extreme compression or `--upload-ratelimit` pacing plus inotify/SIGSTOP hacks, with maintainer warning such boxes are unsuitable - [Source](https://github.com/borgbackup/borg/issues/7107)
- Single-core guidance converges: `GOMAXPROCS=1` (restic), `-C none|laz4` (borg), `--compression=s2-default|none` + `--max-parallel-file-reads=1` (kopia) minimize CPU/RAM at cost of ratio/throughput.
### Inferences
- Low-RAM behavior is fail-slow/fail-OOM rather than graceful: operators must pre-lower parallelism and compression before the run.
- s2-family (kopia) and lz4/zstd,-4 (borg) are the single-core-friendly choices; zlib/lzma/zstd-best are ratio-first and can dominate a weak CPU.
### Gaps
- No citable SQLite/index RAM formula (bytes-per-file or per-GB-repo) found for current restic/borg/kopia versions; vendor docs describe symptoms and knobs, not a sizing equation.
## What systemd-level controls (MemoryMax, CPUQuota, IOWeight) do shipped units use, and what do docs recommend for background backup daemons?
### Takeaway
Upstream backup tools ship no restrictive resource-control units; all concrete CPU/memory/I/O caps come from downstream/community units and generic systemd resource-control docs, with `Nice=` + `CPUQuota`/`MemoryMax`/`IOWeight` as the recommended trio.
### Cited Findings
- systemd `CPUQuota=` sets a hard ceiling as % of one CPU (100%=1 core, 200%=2 cores) via `cpu.max`/`cpu.cfs_quota_us`; `CPUWeight=` (1–10000, default 100) is only relative under contention - [Source](https://manpages.debian.org/bullseye/systemd/systemd.resource-control.5.en.html); same semantics in Arch man - [Source](https://man.archlinux.org/man/systemd.resource-control.5)
- systemd memory knobs: `MemoryHigh=` soft throttle/reclaim, `MemoryMax=` hard OOM-kill limit (K/M/G/T or % of RAM, `infinity` to disable), `MemorySwapMax=` swap cap; Red Hat recommends `MemoryHigh` as main control, `MemoryMax` as last defense - [Source](https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/8/html/managing_monitoring_and_updating_the_kernel/assembly_configuring-resource-management-using-systemd_managing-monitoring-and-updating-the-kernel)
- Community restic timer units use only `Nice=17` (low CPU priority) plus sandboxing (`ProtectSystem=full`, `PrivateTmp=true`, `NoNewPrivileges=yes`, `RestrictAddressFamilies=`, `SystemCallFilter=`, `AmbientCapabilities=CAP_DAC_READ_SEARCH`), with `RandomizedDelaySec=300` + `Persistent=yes` on the timer - [Source](https://www.wildtechgarden.ca/onepagers/real-life-systemd-timers/)
- Community segmented-borg systemd design sets `CPUQuota=80%` and `MemoryMax=2G` on the backup service template - [Source](https://github.com/JoZapf/segmented-borg-backup-system/blob/refs/heads/main/docs/SYSTEMD.md)
- Modern Debian guidance: background CPU → `nice -n 19`; background disk → `ionice -c 3` (BFQ only); hard ceilings/group fairness → unit with `CPUQuota=`/`MemoryMax=`/`IOWeight=`/`IOReadBandwidthMax=`/`IOWriteBandwidthMax=` or a shared `backup.slice`; one-shots via `systemd-run --scope -p CPUQuota=50% -p MemoryMax=1G -p IOWeight=10` - [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
- Same guide warns `ionice` is silently ignored on default `mq-deadline` SSD schedulers (only BFQ honors classes); `MemoryMax` kills rather than slows (use `MemoryHigh` for pushback); `CPUQuota=100%` means one core - [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
- Example backup slice caps restic at 80 MB/s read / 30 MB/s write via `IOReadBandwidthMax=/dev/sda 80M` + `IOWriteBandwidthMax=/dev/sda 30M` with `IOWeight=10` so foreground pools keep headroom - [Source](https://www.bigiron.cc/guides/process-priorities-nice-ionice-cgroup-quotas)
- No shipped upstream restic/borg/kopia unit found with `MemoryMax`/`CPUQuota`/`IOWeight` preset; ArchWiki/packaging ships timers/services without resource caps (report-writer: treat absence as finding, not oversight).
### Inferences
- Recommended Linux server pattern: `Nice=17–19` + optional `IOSchedulingClass=idle` for soft priority, plus a `backup.slice` with `CPUQuota=50–80%`, `MemoryHigh/Max=1–4G`, `IOWeight=10` and per-device `IOWriteBandwidthMax` for hard isolation.
- `nice`/`ionice` alone are insufficient for noisy-neighbor backups on cgroup-v2/BFQ-mixed fleets; systemd slice quotas are the enforceable layer.
### Gaps
- Could not locate a distribution-shipped (Debian/Fedora/Arch package) backup unit that presets `MemoryMax`/`CPUQuota`/`IOWeight`; all cited values are community/blog recommendations, not vendor defaults.
- No citable official restic/borg/kopia doc page prescribing a canonical systemd resource-control stanza as of 2026.