GuestOps/docs/incident-exercise.md
wolf-demon 4dd33c90e1
Some checks are pending
Build and verify web migration / verify (push) Waiting to run
milestone 19 &20
2026-09-29 20:50:40 +01:00

3.0 KiB

Incident and rollback exercise

Run this exercise on the approved sandbox release with synthetic records only. Agree detection, containment and recovery targets before starting. The operator, incident commander and independent reviewer must be different people. Keep actual guest addresses, credentials, tokens, provider payloads and raw incident evidence in the restricted acceptance store.

Gate B scenarios

Record ID Exercise Passing evidence
alert-and-escalate Introduce an approved synthetic failure and use normal monitoring and support routes. The alert is detected, the incident commander is engaged and timestamps meet the agreed detection target.
disable-worker-writes Use the documented controls to stop worker-driven provider writes and FAQ automation. Pending work does not reach a provider, controls remain effective across a worker restart and staff can identify the safe state.
google-uncertain-send Simulate an interrupted or ambiguous Gmail submission. Staff do not resend blindly, reconcile using the stable message ID and retain the disposition.
restore-readiness Use the isolated restore procedure and inspect provider-facing records before enabling workers. The restored system becomes ready, older approvals remain held and external effects are reconciled.
image-rollback Deploy the retained prior image IDs using the rollback procedure without replacing MongoDB or key volumes. The recorded images run, readiness and read-only smoke checks pass and persistent state remains available.
controlled-recovery Recover the approved release after containment. Monitoring is healthy, held operations are dispositioned and the exercise ends with writes disabled and FAQ mode off.

For Gate C, also exercise pms-ambiguous-write and payment-ambiguous-create. In both cases, lose or interrupt the synthetic provider response and prove the operation is reconciled read-only without creating a replacement request or replaying the approval.

Record and validation

Copy deploy/incident-exercise.example.json into the restricted acceptance store. Bind it to the same full release commit and release-record SHA-256 used by the deployment and pilot decision. Replace the example timestamps, people, targets and evidence references. Change a scenario to pass only after its evidence has been independently reviewed. The example is intentionally invalid because its scenarios are not-run.

For a Gate C exercise, set targetGate to C and add the PMS and payment scenario records. Validate the completed record with:

python3 deploy/incident_exercise.py /secure/acceptance/incident-exercise.json

The validator checks record completeness, release binding, scenario coverage, independent roles, timing targets and the safe final state. It does not inspect the referenced evidence, create an incident response capability or authorize a pilot. Retain its output and the record checksum, then reference them from incident-support in the pilot approval record.