Internal. The customer-facing half of this is one paragraph in
Operations.
The model
Git is the source of record. The bucket is a distribution copy.
Worst case — the bucket is gone entirely — is hours of republish, zero data
loss, and no customer action. Every artifact in it was built from a tagged
release and can be built again; what could not be rebuilt is what was
published, at which sequence, to whom — and that lives in
vexa-stations,
where channel.yaml is the authority for entry_seq and the copy inside each
published entry is derived from it.
A pinned station keeps running through all of it. Its Argo CD is reconciling
against digests it already pulled; an unreachable channel stops upgrades, not
workloads. A station following * stops seeing new entries until the channel
is back, which is the same as us not publishing.
What protects the bucket
Versioning is on. An overwrite or a delete does not destroy the previous bytes —list-object-versions still returns them, and a deleted key comes back
from its version id. This is the protection against the ordinary accident (a
wrong oras push, a mistyped s3 rm), which is the one that actually happens.
A nightly replica runs on the replica host at 04:20 host time
(/home/dima/bin/vexa-channel-replica.sh, in crontab), landing in
<replica-path>/data with last-run.json naming the object
count and total size of the run. It uses rclone sync --backup-dir, so an
object that vanishes upstream is moved into deleted/<date>/ rather than
dropped: an accidental mass-delete upstream cannot replicate itself into a
loss here. The script decrypts its own credentials from the secrets vault at run
time into environment variables; nothing credential-shaped is written to disk
on that host.
Station receipts
Everything a customer sent us — the bundle verbatim, its ingest receipt, its manifest — is committed tovexa-stations under
channels/<channel>/stations/<station>/receipts/<timestamp>/. The working copy
in stations/<name>/ on an operator’s laptop is a scratch directory and is
gitignored on purpose; it is not the durable copy and never was.
Keys
The channel signing keys are SOPS+age encrypted in the operator’s secrets vault. Losing the private key is the one failure this design does not absorb: entries and images already signed keep verifying against the published public key, but nothing new can be signed into that channel, and every subscriber has pinned that key. Key custody is therefore a separate discipline from bucket durability, and the age identity that decrypts the vault must exist in more than one place.Rebuilding from nothing
- Read
channels/<channel>/channel.yamlfor the lastentry_seqand the release each entry named. - Re-run the crank in RUNBOOK §1 per release, or restore straight from the off-site replica if it is intact.
- Re-push the kit and the charts from their tags.
- Move
currentto the newest entry. A*station picks it up on its next poll; a pinned station never noticed.