GuestOps/docs/operations.md

7.8 KiB

Operations and recovery

The owner-only Workspace health page reports database reachability, the worker's last heartbeat, mailbox recovery counts and reply exceptions. A heartbeat older than three minutes is marked stale. It proves that the worker process recently reached MongoDB, not that every provider or job succeeded. Reply totals cover at most the latest 500 conversations in this hotel. The page does not verify backups. /health/ready returns only readiness status and HTTP 503 when the database cannot be reached.

Deployment preflight

Run from the dedicated GuestOps checkout on Linux with Python 3.11 or later, Docker with Compose, and GnuPG installed. The tool supports the supplied three-service Compose deployment and the guestops database only. Docker access is administrator-equivalent; use the designated server operator account. Load the reviewed API, worker and MongoDB images first, configure .env with mode 600, and follow deployment.

python3 deploy/ops.py preflight --offline
# After Nginx, DNS, TLS and services are running:
python3 deploy/ops.py preflight

Preflight checks production settings, shared persistent keys, unpublished MongoDB/worker ports, a loopback API port, required images and mounts, private secret files, and at least 8 GiB free disk. The online check also verifies the configured HTTPS readiness URL. This is not a firewall, capacity or provider acceptance test. Allow additional disk space for the database, images and temporary backup/restore files; the supplied server initially had 18 GiB free.

Encrypted backup

Keep the recovery private key on an administrator-controlled recovery machine, with a securely stored passphrase and a second protected recovery copy. Import only its public key on the server and verify its full 40-character fingerprint through a trusted channel. The tool selects that exact fingerprint; it does not establish who owns the key. Do not put private keys, decrypted backups or provider secrets in GitHub.

gpg --import /secure/path/recovery-public.asc
gpg --fingerprint
install -d -m 700 "$HOME/guestops-backups"
read -r -p 'Verified recovery key fingerprint: ' BACKUP_RECIPIENT
python3 deploy/ops.py backup --recipient "$BACKUP_RECIPIENT" \
  --output "$HOME/guestops-backups/guestops-$(date -u +%Y%m%dT%H%M%SZ).tar.gpg" \
  --confirm-maintenance
unset BACKUP_RECIPIENT

This command causes a maintenance interruption. It stops the API and worker, temporarily pauses MongoDB's TTL expiry monitor, dumps MongoDB, and saves the shared Data Protection keys and deployed configuration. The configuration includes credentials and provider files, so the entire bundle is encrypted. Services and the previous TTL setting are restored in a finally block before encryption finishes. A successful backup message does not prove that restarted services are healthy; run the online preflight afterwards.

No other application or administrator may write to this database during the snapshot. Standalone MongoDB dumps need coordinated writes for consistency; see the MongoDB backup guidance. The tool is intentionally limited to the dedicated stack. It requires the hexadecimal MongoDB root password generated in the deployment guide.

Copy the encrypted file off-server to restricted storage after every successful backup. Retain the exact API, worker and MongoDB images with the release: an image tag alone can change, and the drill requires matching image IDs. Agree the backup schedule, retention and tolerated data loss before a hotel pilot. Scheduling, off-server transfer, deletion and alerting are operator responsibilities in this milestone; none is silently installed.

Temporary plaintext files are held in private directories and removed on normal completion or exceptions. Process termination or power loss can leave temporary data, stopped services or TTL expiry disabled. After an interrupted run, inspect the dedicated project and remove only its identified abandoned temporary directory after securing any recovery material. Restore the recorded TTL setting (normally true) and restart the services:

docker compose exec -T mongo mongosh --quiet --nodb --eval 'const c=new Mongo("mongodb://127.0.0.1"); c.getDB("admin").auth(process.env.MONGO_INITDB_ROOT_USERNAME,process.env.MONGO_INITDB_ROOT_PASSWORD); const r=c.getDB("admin").runCommand({setParameter:1,ttlMonitorEnabled:true}); if(!r.ok)quit(1);'
docker compose start api worker
python3 deploy/ops.py preflight

Use the previous value instead of true if TTL expiry was deliberately disabled beforehand. Do not remove production volumes to recover from a failed backup.

Isolated restore drill

Use a trusted encrypted backup and the recovery key on a Linux recovery machine. Unlock the key through the local GPG agent before running the batch command; never pass its passphrase on the command line. Load the exact reviewed images from the backed-up release first.

python3 deploy/ops.py restore-drill /secure/path/guestops-backup.tar.gpg \
  --api-image guestops-api:REVIEWED_RELEASE \
  --mongo-image mongo:8.0

The drill decrypts into a private temporary directory, checks the four expected files and their hashes, rejects unsafe archive entries, and verifies that the restored keys decrypt a protected test value. It then restores MongoDB into a newly created container with no external network or published ports and compares every restored collection's count and indexes. TTL expiry is disabled in this disposable database so expired token records do not disappear during comparison. The exact temporary container and its volumes are removed afterwards. No worker or production application is started.

The decrypted bundle is capped at 8 GiB for this pilot tool. Leave room for the encrypted file, decrypted archive, extracted files and restored database. GPG's output limit bounds decrypted output; checksums detect corruption, not the identity of the backup creator. Use backups from the trusted operator only. A drill checks counts, indexes and key decryption, not every document's business meaning or live provider connectivity.

Actual disaster recovery

This milestone implements a rehearsal, not an automatic production restore or cutover command. Before real guest data is accepted, rehearse a separate recovery deployment and document its precise image IDs, volume names, secret locations and operator responsibilities.

Restore into new isolated MongoDB and key volumes; preserve the damaged original for investigation. Reconstruct reviewed configuration from the encrypted bundle without copying old Docker container IDs or blindly executing archived configuration. Restore the guestops database with the compatible MongoDB tools and restore the matching key files with the application's required ownership. Check hotel/account records and saved settings before exposing the recovered API. Keep the worker stopped, provider writes disabled and external network access restricted throughout this process.

A restored database can predate emails, invoices and PMS changes that providers already completed. Review pending, sending and uncertain records against provider evidence before enabling any worker, including automatic FAQ rules. Do not replay an older approval merely because the restored record says it is pending. Reconcile external effects, validate account sessions and mailbox authorization, and explicitly approve the cutover only after these checks. Rotate credentials if compromise prompted the recovery. Keep the old deployment stopped when enabling the replacement.

CI exercises a synthetic encrypted backup and isolated restore drill, including actual key decryption and database comparison. A successful CI drill is separate from the required rehearsal on the Debian server with its actual deployment configuration.