@@ -0,0 +1,54 @@
|
||||
---
|
||||
id: "OPS.INC.001"
|
||||
category: "operations/incidents"
|
||||
revision-version: 1
|
||||
revision-date: "2024-01-14"
|
||||
estimate-time: "30m"
|
||||
necessity: "mandatory"
|
||||
frameworks:
|
||||
- name: "soc2"
|
||||
sections: ["CC2.2"]
|
||||
---
|
||||
|
||||
# Build an incident response process
|
||||
|
||||
## Purpose
|
||||
|
||||
An incident response plan is crucial for quickly identifying, containing, and
|
||||
resolving incidents, minimizing potential disruptions. It ensures you’re
|
||||
prepared effectively and that you keep your operations running smoothly, even in
|
||||
the face of unexpected threats.
|
||||
|
||||
## Implementation
|
||||
|
||||
### Build an Incident Response Plan and share it with your team.
|
||||
|
||||
Template:
|
||||
|
||||
[Incident response plan](templates/incident_response_plan.md)
|
||||
|
||||
### Create a slack channel
|
||||
|
||||
Create a slack channel to allow your team to raise incident and be kept aware of
|
||||
their resolution
|
||||
|
||||
1. Create a channel (eg `#incident`)
|
||||
2. Pin the **Incident Response Plan** in the channel for quick access.
|
||||
3. Establish guidelines for using the channel, including:
|
||||
1. How to report incidents (e.g., “Post a message with a brief summary and
|
||||
relevant context”).
|
||||
2. Expected response times for acknowledgment and updates.
|
||||
4. Assign roles for incident management, such as:
|
||||
1. **On-duty engineer**: Monitors and triages notifications.
|
||||
2. **Response team**: Engineers responsible for handling and resolving
|
||||
incidents.
|
||||
5. Set up reminders or periodic messages in Slack to encourage proactive
|
||||
monitoring and review of unresolved issues.
|
||||
6. Integrate relevant notifications, if applicable, into the Slack channel for
|
||||
visibility (e.g., system alerts or logs).
|
||||
|
||||
## Evidence
|
||||
|
||||
- Share your incident response plan
|
||||
- screenshot of your #incident slack channel being used (with postmortem
|
||||
visible)
|
||||
45
controls/operations/incidents/OPS.INC.002_drp.md
Normal file
45
controls/operations/incidents/OPS.INC.002_drp.md
Normal file
@@ -0,0 +1,45 @@
|
||||
---
|
||||
id: "OPS.INC.002"
|
||||
category: "operations/incidents"
|
||||
revision-version: 1
|
||||
revision-date: "2024-01-14"
|
||||
estimate-time: "6h"
|
||||
necessity: "mandatory"
|
||||
frameworks:
|
||||
- name: "soc2"
|
||||
sections: ["CC2.2"]
|
||||
---
|
||||
|
||||
# Disaster Recovery Plan (DRP)
|
||||
|
||||
## Purpose
|
||||
|
||||
Disasters might seem far-fetched, but data loss, service outages, and
|
||||
misconfigurations can happen. Whether it's a simple human error or a regional
|
||||
outage, a DRP helps your team recover fast and maintain customer trust.
|
||||
|
||||
## Implementation
|
||||
|
||||
Here is a basic template, to use as a good starting point:
|
||||
|
||||
[DRP Template](https://www.notion.so/DRP-Template-13e1cc0bd5bc800d9db2f77d3c884521?pvs=21)
|
||||
|
||||
<aside>
|
||||
💡
|
||||
|
||||
If it is easier to manage, you can have several DRP: one for your database, one
|
||||
for your network, one for your infra, etc… that way, when testing, you can test
|
||||
smaller part of it and it adapts better to potential real life events.
|
||||
|
||||
</aside>
|
||||
|
||||
The idea is to be prepared for the unexpected. You don’t need to overthink
|
||||
things, it’s more to know what to do when you will be in the rush of a regional
|
||||
outage. For that, your DRP needs to cover:
|
||||
|
||||
1. **Scope & timings**: clear scope and the expected time to do it
|
||||
2. **Step-by-step:** guide on how to proceed
|
||||
3. **Contacts**: list of internal and external contacts crucial during a
|
||||
disaster.
|
||||
4. **Testing and Review Log**: A record of tests performed to validate the DRP’s
|
||||
effectiveness and any improvements made.
|
||||
@@ -0,0 +1,50 @@
|
||||
# DRP Template
|
||||
|
||||
### **Overview and Objectives**
|
||||
|
||||
- **Purpose**: Provide step-by-step instructions to recover **[Critical
|
||||
Services]** in case of a disaster, ensuring minimum downtime and data
|
||||
protection.
|
||||
- **Scope**: This plan applies to **[Specify Services, e.g., Cloud SQL,
|
||||
application servers, etc.]**.
|
||||
- **Objectives**: Restore critical services within **[Define Recovery Time
|
||||
Objective, e.g., 2 hours]**.
|
||||
|
||||
### Instructions
|
||||
|
||||
<aside>
|
||||
⚠️
|
||||
|
||||
- Don’t overthink it.
|
||||
- Unless specifically required by a customer, don’t change your cloud provider
|
||||
for the recovery
|
||||
</aside>
|
||||
|
||||
1. Create a VM
|
||||
2. Run the following commands to log-in and to install the minimum software
|
||||
|
||||
`your commands`
|
||||
|
||||
3. Allow the VM to connect to the Database
|
||||
|
||||
`your commands`
|
||||
|
||||
4. Check the service is running
|
||||
|
||||
### **Troubleshooting**
|
||||
|
||||
> In this section provide everything you think it can help people running the
|
||||
> instruction e.g. “Connection Refused”, “Docker Image does not exist”, etc.
|
||||
|
||||
### **Contact Information**
|
||||
|
||||
| Role | Name | Contact Info | Backup Contact |
|
||||
| ---------------------- | ----------- | ---------------------- | --------------------- |
|
||||
| CTO | [Name] | [Email, Phone] | [Backup Contact] |
|
||||
| Cloud Provider Support | GCP Support | [Support Contact Info] | [Alternative Contact] |
|
||||
|
||||
### **Testing and Review Log**
|
||||
|
||||
| Date | Test Type | Result | Notes/Improvements |
|
||||
| ------------ | ----------------------- | ---------- | ------------------------------------- |
|
||||
| [YYYY-MM-DD] | Backup Restoration Test | Successful | Updated recovery steps for [Service]. |
|
||||
@@ -0,0 +1,63 @@
|
||||
### **1. Triage & validate**
|
||||
|
||||
The on-duty engineer receives the notification, he/she analyze logs, system
|
||||
behaviors, and context related to the indication versus known behaviors.
|
||||
|
||||
He/She needs to determine whether the notification is legit or not:
|
||||
|
||||
1. If it is legit ⇒ it is promoted to an ongoing incident. In the thread, CTO is
|
||||
notified and a message is posted with:
|
||||
|
||||
- A brief summary of the incident
|
||||
- Any immediate risks or concerns
|
||||
- Preliminary findings and scope (if known)
|
||||
- Ping for the relevant engineer (the response team).
|
||||
|
||||
⇒ the response team commences work.
|
||||
|
||||
2. If it is not (ie a false alarm), a message is posted in the thread with the
|
||||
reason for dismissal. The on-duty engineer needs to:
|
||||
- Determine and address the root cause for the false alarm. He/she needs to
|
||||
remediate any system or process issues that led to it.
|
||||
- End the incident response process
|
||||
|
||||
### **2. Handling & Resolution**
|
||||
|
||||
Resolution occurs in 3 steps detailed below. All decisions or actions taken must
|
||||
be documented.
|
||||
|
||||
1. **Contain:**
|
||||
|
||||
If possible, isolate system to prevent the incident from spreading or causing
|
||||
additional damage
|
||||
|
||||
1. **Eradicate:**
|
||||
|
||||
Identify and remove the root cause of the incident - roll back is a strong
|
||||
option.
|
||||
|
||||
1. **Recover:**
|
||||
|
||||
Resume normal business operations by restoring system functionality. Keep strong
|
||||
monitoring to catch any signs of repetition.
|
||||
|
||||
At the end of this process, an update is posted on the thread to keep all team
|
||||
informed. If the complete process takes longer than 30min, a quick update will
|
||||
be given in the thread every 30min so stakeholders are aware of the current
|
||||
state.
|
||||
|
||||
### **3. Postmortem**
|
||||
|
||||
Once the incident is resolved, the on-duty engineer document the findings,
|
||||
decisions and actions taken during the event in a document to which he/she joins
|
||||
recommendations. This document must be shared with stakeholders.
|
||||
|
||||
The postmortem document must answer:
|
||||
|
||||
- What was the incident? What happened?
|
||||
- Which decisions or actions were taken to respond?
|
||||
- What could have prevented it from happening?
|
||||
- What can be improved in terms of security, processes, communication, tools,
|
||||
etc.?
|
||||
|
||||
The CTO must validate and ensure that recommended improvements are implemented.
|
||||
Reference in New Issue
Block a user