Stabilise the service, coordinate decisions, preserve evidence, and keep stakeholders informed during an incident.
- incident response
- production incident
- incident management
- service recovery
Establish one incident lead
Assign a coordinator, technical responders, and a communications owner. Use one shared timeline so people do not make conflicting changes or duplicate investigations.
Cloud, Data & Security
Thoughtful decisions compound over time.
Practical product work brings technical choices back to the people and workflows they are meant to serve.
Limit impact before chasing every cause
Confirm affected systems and users, consider safe rollback or containment, and preserve useful logs. Avoid destructive actions that could remove evidence or make recovery harder.
Communicate and learn
Provide concise updates at agreed intervals and record decisions. After recovery, run a blameless review that identifies practical improvements to detection and response.
Practical application
Open one incident channel, appoint a coordinator, record a timeline, and identify the customer-facing impact. Prefer a reversible mitigation, communicate an update cadence, and preserve evidence before making destructive changes.