# Operations and recovery The owner-only **Workspace health** page reports database reachability, the worker's last heartbeat, mailbox recovery counts and reply exceptions. A heartbeat older than three minutes is marked stale. It proves that the worker process recently reached MongoDB, not that every provider or job succeeded. Reply totals cover at most the latest 500 conversations in this hotel. The page does not verify backups. `/health/ready` returns only readiness status and HTTP 503 when the database cannot be reached. ## Release evidence and rollback Every non-pull-request CI build packages the API and worker images under the full Git commit SHA. The accompanying `release-record.json` binds the archive checksum, application version, commit, image references and immutable Docker image IDs. Retain both files together in restricted off-host release storage; the CI artifact is a transfer mechanism, not the permanent archive. Before deployment, verify the archive against its record without loading it: ```sh python3 - <<'PY' import hashlib, json, pathlib r = json.load(open('release-record.json', encoding='utf-8')) p = pathlib.Path(r['artifact']['name']) assert hashlib.sha256(p.read_bytes()).hexdigest() == r['artifact']['sha256'] print(r['commit'], r['version'], r['images']) PY ``` Load the archive, verify each loaded image ID matches the record, set `GUESTOPS_API_IMAGE` and `GUESTOPS_WORKER_IMAGE` to the recorded full-SHA references, and run the deployment preflight. Record the CI run, commit, checksum and operator in the change ticket. A release tag is an approval marker; do not move or reuse an existing tag. The application and web versions must match before the record can be created. For rollback, first disable worker-driven external writes and reconcile any sending, payment or PMS operation that may have completed since the prior release. Confirm the previous release archive and record are retained, verify its checksum and image IDs, take an encrypted backup, then select the previous recorded image references in `.env` and recreate only the API and worker. Do not roll back MongoDB or the key volume merely to change application images. Run the online preflight, readiness check and read-only smoke test before re-enabling the worker or provider writes. If a release introduced an incompatible data change, follow its release-specific recovery plan rather than starting an older image against newer data. ## Deployment preflight Run from the dedicated GuestOps checkout on Linux with Python 3.11 or later, Docker with Compose, and GnuPG installed. The tool supports the supplied three-service Compose deployment and the `guestops` database only. Docker access is administrator-equivalent; use the designated server operator account. Load the reviewed API, worker and MongoDB images first, configure `.env` with mode 600, and follow [deployment](deployment.md). ```sh python3 deploy/ops.py preflight --offline # After Nginx, DNS, TLS and services are running: python3 deploy/ops.py preflight ``` Preflight checks production settings, shared persistent keys, unpublished MongoDB/worker ports, a loopback API port, required images and mounts, private secret files, and at least 8 GiB free disk. The online check also verifies the configured HTTPS readiness URL. This is not a firewall, capacity or provider acceptance test. Allow additional disk space for the database, images and temporary backup/restore files; the supplied server initially had 18 GiB free. Configured provider files must be readable by the API container's `app` user while inaccessible to other users. Set their ownership to that image's application UID and mode 600, and run the backup as an authorized operator able to read them (for example root with its designated public GPG keyring). Keep their parent directory private. Empty shipped example files contain no credentials and do not need this ownership change. Check the UID in the reviewed image rather than assuming it matches your host login. ## Encrypted backup Keep the recovery private key on an administrator-controlled recovery machine, with a securely stored passphrase and a second protected recovery copy. Import only its public key on the server and verify its full 40-character fingerprint through a trusted channel. The tool selects that exact fingerprint; it does not establish who owns the key. Do not put private keys, decrypted backups or provider secrets in GitHub. ```sh gpg --import /secure/path/recovery-public.asc gpg --fingerprint install -d -m 700 "$HOME/guestops-backups" read -r -p 'Verified recovery key fingerprint: ' BACKUP_RECIPIENT python3 deploy/ops.py backup --recipient "$BACKUP_RECIPIENT" \ --output "$HOME/guestops-backups/guestops-$(date -u +%Y%m%dT%H%M%SZ).tar.gpg" \ --confirm-maintenance unset BACKUP_RECIPIENT ``` This command causes a maintenance interruption. It stops the API and worker, temporarily pauses MongoDB's TTL expiry monitor, dumps MongoDB, and saves the shared Data Protection keys and deployed configuration. The configuration includes credentials and provider files, so the entire bundle is encrypted. Services and the previous TTL setting are restored in a `finally` block before encryption finishes. A successful backup message does not prove that restarted services are healthy; run the online preflight afterwards. No other application or administrator may write to this database during the snapshot. Standalone MongoDB dumps need coordinated writes for consistency; see the [MongoDB backup guidance](https://www.mongodb.com/docs/v8.0/tutorial/backup-and-restore-tools/). The tool is intentionally limited to the dedicated stack. It requires the hexadecimal MongoDB root password generated in the deployment guide. Copy the encrypted file off-server to restricted storage after every successful backup. Retain the exact API, worker and MongoDB images with the release: an image tag alone can change, and the drill requires matching image IDs. Agree the backup schedule, retention and tolerated data loss before a hotel pilot. Scheduling, off-server transfer, deletion and alerting are operator responsibilities in this milestone; none is silently installed. ### Optional systemd schedule The repository includes an opt-in daily systemd service and timer. They are templates, not automatically installed. The service assumes the reviewed checkout is `/srv/guestops` and stages encrypted files in `/var/backups/guestops`; review and change both unit files together if the host uses different paths. Create the staging directory and environment file without storing a private key or passphrase on the server: ```sh sudo install -d -m 700 /var/backups/guestops /var/lib/guestops-backup/gnupg /etc/guestops sudo install -m 600 /dev/null /etc/guestops/backup.env sudoedit /etc/guestops/backup.env sudo env GNUPGHOME=/var/lib/guestops-backup/gnupg gpg --import /secure/path/recovery-public.asc sudo env GNUPGHOME=/var/lib/guestops-backup/gnupg gpg --fingerprint ``` The file contains only these two settings. `BACKUP_RECIPIENT` is the verified 40-character public-key fingerprint already imported into the service account's GPG keyring. ```text BACKUP_RECIPIENT=0123456789ABCDEF0123456789ABCDEF01234567 BACKUP_DIRECTORY=/var/backups/guestops ``` Review the unit sandbox against the host, copy `deploy/systemd/guestops-backup.service` and `.timer` to `/etc/systemd/system`, then validate and test before enabling: ```sh sudo systemd-analyze verify /etc/systemd/system/guestops-backup.service /etc/systemd/system/guestops-backup.timer sudo systemctl daemon-reload sudo systemctl start guestops-backup.service sudo systemctl status guestops-backup.service sudo systemctl enable --now guestops-backup.timer systemctl list-timers guestops-backup.timer ``` The timer deliberately causes the same brief maintenance interruption as a manual backup. `Persistent=true` runs a missed event after downtime, so choose and communicate the maintenance window. A successful unit only stages an encrypted file locally. Configure an independently monitored off-host transfer, verify the destination checksum, alert on both unit and transfer failure, and test the alert route. Do not add automatic deletion until retention, legal hold and recovery requirements have named owners. Temporary plaintext files are held in private directories and removed on normal completion or exceptions. Process termination or power loss can leave temporary data, stopped services or TTL expiry disabled. After an interrupted run, inspect the dedicated project and remove only its identified abandoned temporary directory after securing any recovery material. Restore the recorded TTL setting (normally true) and restart the services: ```sh docker compose exec -T mongo mongosh --quiet --nodb --eval 'const c=new Mongo("mongodb://127.0.0.1"); c.getDB("admin").auth(process.env.MONGO_INITDB_ROOT_USERNAME,process.env.MONGO_INITDB_ROOT_PASSWORD); const r=c.getDB("admin").runCommand({setParameter:1,ttlMonitorEnabled:true}); if(!r.ok)quit(1);' docker compose start api worker python3 deploy/ops.py preflight ``` Use the previous value instead of true if TTL expiry was deliberately disabled beforehand. Do not remove production volumes to recover from a failed backup. ## Isolated restore drill Use a trusted encrypted backup and the recovery key on a Linux recovery machine. Unlock the key through the local GPG agent before running the batch command; never pass its passphrase on the command line. Load the exact reviewed images from the backed-up release first. ```sh python3 deploy/ops.py restore-drill /secure/path/guestops-backup.tar.gpg \ --api-image guestops-api:REVIEWED_RELEASE \ --mongo-image mongo:8.0 ``` The drill decrypts into a private temporary directory, checks the four expected files and their hashes, rejects unsafe archive entries, and verifies that the restored keys decrypt a protected test value. It then restores MongoDB into a newly created container with no external network or published ports and compares every restored collection's count and indexes. TTL expiry is disabled in this disposable database so expired token records do not disappear during comparison. The exact temporary container and its volumes are removed afterwards. No worker or production application is started. The decrypted bundle is capped at 8 GiB for this pilot tool. Leave room for the encrypted file, decrypted archive, extracted files and restored database. GPG's [output limit](https://www.gnupg.org/documentation/manuals/gnupg/GPG-Input-and-Output.html) bounds decrypted output; checksums detect corruption, not the identity of the backup creator. Use backups from the trusted operator only. A drill checks counts, indexes and key decryption, not every document's business meaning or live provider connectivity. ## Actual disaster recovery This milestone implements a rehearsal, not an automatic production restore or cutover command. Before real guest data is accepted, rehearse a separate recovery deployment and document its precise image IDs, volume names, secret locations and operator responsibilities. Restore into new isolated MongoDB and key volumes; preserve the damaged original for investigation. Reconstruct reviewed configuration from the encrypted bundle without copying old Docker container IDs or blindly executing archived configuration. Restore the `guestops` database with the compatible MongoDB tools and restore the matching key files with the application's required ownership. Check hotel/account records and saved settings before exposing the recovered API. Keep the worker stopped, provider writes disabled and external network access restricted throughout this process. **A restored database can predate emails, invoices and PMS changes that providers already completed.** Review pending, sending and uncertain records against provider evidence before enabling any worker, including automatic FAQ rules. Do not replay an older approval merely because the restored record says it is pending. Reconcile external effects, validate account sessions and mailbox authorization, and explicitly approve the cutover only after these checks. Rotate credentials if compromise prompted the recovery. Keep the old deployment stopped when enabling the replacement. CI exercises a synthetic encrypted backup and isolated restore drill, including actual key decryption and database comparison. A successful CI drill is separate from the required rehearsal on the Debian server with its actual deployment configuration.