Major Incident Response Process
ADME / OSDU Platform — Equinor
| Document Owner | OSDU Platform Team |
| Version | 1.0 |
| Created | August 2026 |
| Based On | DR Test May 2026 — Lessons Learnt & DR Trigger & Response Process |
| Status | Draft |
Purpose
This document defines the end-to-end process for detecting, declaring, and responding to a major incident on the Equinor ADME / OSDU platform. It applies across all environments — development, test, and production — and is intended to provide a consistent, repeatable response framework regardless of the specific incident type.
A major incident is any event that causes, or has the potential to cause, significant disruption to platform availability, data integrity, or the ability of data producers and consuming applications to operate. A Disaster Recovery (DR) restore is one example of a major incident; other examples include large-scale data corruption, platform-wide service outages, and confirmed security breaches.
This document does not replace the DR-specific runbook. For DR restore steps on the production instance, refer to the DR Trigger & Response Process document.
Scope
Environments Covered
| Environment | Examples | Incident Handling |
|---|---|---|
Production (data) |
Live platform serving all data producers and consuming applications | Full incident process applies; highest urgency |
Test (test, staging) |
Pre-production validation, DR rehearsals, integration testing | Full incident process applies where data integrity or test fidelity is at risk; lower urgency than production |
| Development | Developer sandboxes, feature branches, spike environments | Lightweight process applies; team-level resolution preferred before escalation |
Incident Types Covered
The following categories constitute a major incident on this platform:
| Category | Examples |
|---|---|
| Data Loss / Corruption | Malicious or confirmed unauthorised deletion of records, deletion of a data partition, deliberate corruption of ACLs or legal tags |
| Platform Availability | Full or partial unavailability of ADME services (Storage, Search, Indexer, Seismic DMS, Entitlements), data partition unresponsive for more than 30 minutes to 1 hour |
| Security Breach | Compromised service principal credentials, unauthorised access, suspected ransomware or cyberattack, privilege escalation |
| Infrastructure Failure | Azure region/datacenter failure, backing storage (GCS/blob) inaccessible, network connectivity loss to ADME |
| Dependency Failure | Critical upstream or downstream service failure affecting data flow or consuming applications (e.g. SMDA, SDMC, identity provider) |
What This Process Does NOT Cover
- Data corrections that can be addressed by a data producer re-running their pipeline
- Planned maintenance or environment refreshes
- Individual pipeline failures or ingestion errors with no broader platform impact
- Feature bugs or configuration changes in a single service that do not affect data integrity or platform availability
Incident Severity Levels
| Level | Name | Definition | Response Target |
|---|---|---|---|
| P1 — Critical | Production environment severely degraded, significant data loss or corruption confirmed, or active security breach | Immediate response; War Room convened within 30 minutes; 24/7 until resolved | |
| P2 — Major | Production environment partially degraded, or data loss contained to a subset of kinds/partitions, or test environment fully unavailable ahead of a scheduled test | Response within 1 hour; dedicated coordination channel opened | |
| P3 — Significant | Test or dev environment degraded; production impact contained or under investigation; potential data integrity concern | Response within 4 hours during business hours | |
| P4 — Minor | Dev environment issue; no production or test impact; investigation required but no immediate risk | Response within 1 business day |
Environment modifier: Any incident directly impacting the production
datapartition is automatically elevated to P1 or P2 regardless of apparent scope, until confirmed otherwise. If a Microsoft support ticket is required, it must be raised at the same severity level (Severity A / Critical for P1; Severity B for P2) and must not be downgraded without explicit agreement from the OSDU Platform Team Lead.
Detection — How to Identify a Major Incident
Monitoring Signals
The following signals may indicate a major incident. Any team member or data producer who observes these should raise immediately with the OSDU Platform Team:
| Signal | Source | Potential Incident Type |
|---|---|---|
| Mass record deletion or unexpected record count drop | ADME Search API / audit logs | Data Loss / DR Event |
| Bulk ACL or entitlement group deletion | OEP EntitlementsLogs | Data Loss / Security |
| Service returning repeated 500 / 503 errors | ADME health monitoring / Azure Monitor | Platform Availability |
| Seismic subproject inaccessible or missing | Seismic DMS API | Data Loss / Platform Availability |
| Data partition unresponsive or deleted | Azure Portal / ADME health | DR Event / Infrastructure Failure |
| Unusual or off-hours service principal activity | Azure AD sign-in logs | Security Breach |
| Downstream application reporting unexpected missing or inaccessible data | Direct report from application or data producer | Data Loss / Platform Availability |
| Pipeline failures across multiple data producer teams simultaneously | Data producer monitoring | Platform Availability / Infrastructure Failure |
| Azure region or datacenter health alert | Azure Service Health | Infrastructure Failure |
Reporting by Data Producers and Application Managers
⚠️ If you are a data producer or application manager and notice anything unusual — missing records, unexpected query failures, data that should be present but isn't, or behaviour that suggests data may have been lost or corrupted — contact the OSDU Platform Team immediately and mark your message as high urgency. Do not assume the issue is on your side without first checking with the platform team. Early detection is critical.
Recommended Alert Setup
The following should be in place as a minimum on the production instance:
- Azure Monitor Alert — metric alert for a spike in HTTP 4xx/5xx error rates or DELETE operations within a rolling 15-minute window.
- Log Analytics Scheduled Query — runs every 30 minutes against ADME Storage and OEP EntitlementsLogs, flagging bulk DELETEs above a defined threshold.
- Record Count Baseline Check — automated comparison of record counts per OSDU kind against a documented baseline; triggers alert on >10% drop.
- Azure Service Health Alerts — configured to notify the team of Azure region or ADME service health events.
- Seismic Subproject Poll — periodic check confirming subproject presence and GCS bucket accessibility.
Incident Declaration
Who Can Declare
Any OSDU Platform Team member can initiate an incident investigation. Only the OSDU Platform Team Lead (or their named deputy) can formally declare a major incident.
Declaration Criteria
Declare a major incident if any of the following are true:
- Data loss or corruption is confirmed or strongly suspected at scale (see Section 2 for severity)
- A production ADME service is unavailable and the cause is not immediately resolvable
- A security breach or compromise is suspected
- A test environment is non-functional and a scheduled test or milestone is imminent
- Equinor's DR disaster scenarios (Cyber-attack / Physical Failure / Data Deletion) are triggered
When in doubt, declare. It is always better to stand down from an incident than to discover one was not declared in time.
Response Phases
📢 War Room: For P1/P2 incidents, all associated roles will be called into a dedicated War Room (Teams bridge call or channel) immediately on declaration. This ensures all stakeholders are aligned, decisions are made rapidly, and the response is coordinated from a single point.
Detection & Initial Assessment (Target: within 30 minutes)
| Step | Action | Responsible | Time Target |
|---|---|---|---|
| 1.1 | Alert fires, anomaly observed, or report received from data producer / application manager — treat all reports as high-urgency until investigated | Monitoring system / Any team member | T+0 |
| 1.2 | Assign an incident lead to coordinate the initial response | OSDU Platform Team Lead | T+0 to T+10 min |
| 1.3 | Run initial health checks: record counts, service status, audit logs, Azure Portal | OSDU Platform Team | T+0 to T+15 min |
| 1.4 | Determine environment affected (prod / test / dev) and confirm whether incident is real or a false positive | OSDU Platform Team | T+15 min |
| 1.5 | Classify against severity level (P1–P4) and incident type (Section 2) | OSDU Platform Team Lead | T+15 to T+30 min |
| 1.6 | Formally declare major incident (if applicable) and open coordination channel | OSDU Platform Team Lead | T+30 min |
Phase 2: Stakeholder Notification & Triage (Target: within 1 hour of declaration)
| Step | Action | Responsible | Time Target |
|---|---|---|---|
| 2.1 | Open dedicated Teams bridge call / channel; invite all stakeholders per Section 6 | OSDU Platform Team Lead | T+30 min |
| 2.2 | Send initial incident notification (environment, nature of incident, current status) | OSDU Platform Team Lead | T+30 to T+45 min |
| 2.3 | Engage Microsoft via support ticket if ADME platform action is required — classify as Severity A / Critical for P1/P2 | OSDU Platform Team | T+30 to T+45 min |
| 2.4 | Notify Equinor Security Team if a breach or credential compromise is suspected | OSDU Platform Team Lead | T+30 to T+45 min |
| 2.5 | Instruct data producer teams to pause all ingest pipelines to the affected partition — P1/P2 only | OSDU Platform Team | T+45 min to T+1 hr |
| 2.6 | Notify downstream Application Managers of expected impact and outage window | OSDU Platform Team Lead | T+1 hr |
Phase 3: Investigation & Decision (Target: within 2–3 hours of declaration)
| Step | Action | Responsible | Time Target |
|---|---|---|---|
| 3.1 | Investigate root cause — review ADME audit logs, Azure AD sign-in logs, OEP EntitlementsLogs, and Azure Monitor | OSDU Platform Team | T+1 to T+2 hr |
| 3.2 | Determine whether the incident can be resolved through targeted remediation (reingestion, manual fix, config rollback) or requires a full DR restore | OSDU Platform Team Lead | T+2 hr |
| 3.3 | If DR restore is required: follow the DR Trigger & Response Process; obtain Task Lead approval before proceeding | Task Lead + Platform Team Lead | T+2 hr |
| 3.4 | If security breach is suspected: rotate service principal credentials and confirm with Equinor Security before any restore or repair action | Platform Team + Equinor Security | T+2 hr |
| 3.5 | Document the decision: scope, affected data, chosen remediation approach, approving stakeholder | OSDU Platform Team | T+2 hr |
| 3.6 | Confirm remediation plan with Microsoft if their involvement is required | OSDU Platform Team | T+2 to T+3 hr |
Phase 4: Remediation Execution
| Step | Action | Responsible | Time Target |
|---|---|---|---|
| 4.1 | Execute agreed remediation — DR restore, service restart, config rollback, data reingestion, or security remediation | OSDU Platform Team / Microsoft | Per agreed plan |
| 4.2 | Provide regular updates (minimum hourly) to all stakeholders in the coordination channel | OSDU Platform Team Lead | During execution |
| 4.3 | Do not run write or ingest operations against the affected partition during active remediation | All teams | During execution |
| 4.4 | Monitor for unexpected side effects — watch for 429s, 500s, or secondary failures | OSDU Platform Team | During execution |
Phase 5: Validation (Target: within 24 hours of remediation completion)
| Step | Action | Responsible | Time Target |
|---|---|---|---|
| 5.1 | Run platform health checks: service availability, record count checks against baseline, ACL verification | OSDU Platform Team | T+0 to T+2 hr post-remediation |
| 5.2 | Run smoke test queries across affected OSDU data kinds | OSDU Platform Team | T+2 to T+4 hr post-remediation |
| 5.3 | Data producer teams confirm their respective data is accessible and intact | Data Producers | T+4 to T+8 hr post-remediation |
| 5.4 | Application Managers confirm end-to-end application access | Application Managers | T+8 to T+24 hr post-remediation |
| 5.5 | If DR restore was executed: verify seismic subproject and GCS bucket accessibility explicitly | OSDU Platform Team + SDMC | T+2 hr post-restore |
| 5.6 | Confirm recovery with Task Lead and formally close the incident | Task Lead + Platform Team Lead | T+24 hr post-remediation |
Phase 6: Post-Incident (within 5 business days)
| Step | Action | Responsible |
|---|---|---|
| 6.1 | Produce incident report: timeline, root cause, scope of impact, actions taken, data lost (if any) | OSDU Platform Team |
| 6.2 | Conduct retrospective / lessons learnt session with all stakeholders | OSDU Platform Team Lead |
| 6.3 | Update runbooks and this process document with gaps discovered | OSDU Platform Team |
| 6.4 | Raise backlog items in ADO / Jira for follow-up actions | OSDU Platform Team |
| 6.5 | Review and update severity thresholds and escalation paths if required | OSDU Platform Team + Business |
Roles & Responsibilities
| Role | Responsibility | Notification Priority |
|---|---|---|
| OSDU Platform Team Lead | Incident declaration, stakeholder coordination, Microsoft liaison, approval facilitation | Immediate — T+0 |
| OSDU Platform Team (engineers) | Detection, log investigation, runbook execution, health checks, ACL remediation | Immediate — T+0 |
| Task Lead | Formal approval for DR restore or major remediation actions; business impact assessment; incident closure | Within 1 hour of P1/P2 declaration |
| Microsoft (ADME Support) | Restore execution (DR events); root cause investigation; GCS bucket and infrastructure remediation | Within 30–45 min via Severity A ticket (P1/P2) |
| Equinor Security Team | Credential rotation; security incident assessment; breach containment | Within 1 hour if compromise suspected |
| Data Producer Leads | Pause ingest pipelines on instruction; validate respective data post-remediation | Within 1 hour of declaration |
| Application Managers | Pause non-critical operations during remediation; participate in end-to-end validation | Notified within 1 hour; active from Phase 5 |
| App Sec | Security assessment and oversight; advise on credential rotation and breach scope | Within 1 hour if compromise suspected |
| Data Office | Data governance oversight; confirm data integrity and compliance post-incident | Notified within 1 hour of P1/P2 declaration |
Communication Plan
Channels & Cadence
| Scenario | Channel | Frequency |
|---|---|---|
| Initial declaration | Email + Teams message to all stakeholders | Once — immediately on declaration |
| Active remediation | Dedicated Teams bridge call or channel | Hourly updates while remediation is in progress |
| Stakeholder updates (business hours) | Teams channel | Every 2 hours |
| Out-of-hours P1 incident | On-call contact list (to be defined) | As needed; continuous until resolved |
| Microsoft communication | Severity A support ticket + direct Teams/email | Minimum hourly during active engagement |
Communication Templates
Initial Declaration Message:
🔴 [MAJOR INCIDENT DECLARED — P{severity}] Environment: {prod / test / dev} Nature: {brief description of incident type and impact} Status: Under investigation Incident Lead: {name} Bridge / Channel: {link} Next update: {time}
Incident Closure Message:
✅ [INCIDENT CLOSED — P{severity}] Resolution: {brief description of remediation applied} Impact summary: {data affected, downtime, etc.} Incident report: {link — to follow within 5 business days} Retrospective: {date/time — to be scheduled}
Key Contacts (to be populated)
| Role | Name | Contact |
|---|---|---|
| OSDU Platform Team Lead | TBD | TBD |
| Task Lead | TBD | TBD |
| Task Lead Deputy (out-of-hours) | TBD | TBD |
| Microsoft Account Contact | TBD | TBD |
| Microsoft Support (Severity A) | Azure Portal — support ticket | aka.ms/azuresupport |
| Equinor Security On-Call | TBD | TBD |
| Data Producer Leads | TBD | TBD |
| Application Managers | TBD | TBD |
| Data Office | TBD | TBD |
| App Sec | TBD | TBD |
Environment-Specific Considerations
Production (data partition)
- Full incident process applies without exception.
- Any P1/P2 incident must involve Task Lead approval before major remediation actions are taken.
- Attempt targeted remediation before escalating to more disruptive interventions (e.g. full restores, partition-level actions).
- Where Microsoft support is required, raise the ticket before midday CET where possible to avoid overnight delays due to USA-based support. For P1 incidents, Severity A tickets guarantee 24/7 response.
Test Environments (test, staging)
- Full incident process applies where data integrity or test fidelity is at risk (e.g. ahead of a scheduled DR rehearsal or certification activity).
- Test environments are isolated from production — partition isolation must be verified before concluding there is no production impact.
- A simpler, team-level resolution path is acceptable for low-risk test environment issues (P3/P4).
- For planned DR tests, a separate test runbook governs the execution; this process governs any unplanned failures during or after the test.
Development Environments
- Team-level resolution is the default. Escalate to OSDU Platform Team only if the issue cannot be resolved within the team.
- Development environments must not be used as a workaround for production issues.
DR-Specific Considerations
When a major incident results in a DR restore being required, the following apply in addition to this general process:
| Topic | Consideration |
|---|---|
| RPO | ~24 hours — restore runs once daily at midnight UTC. All data written since the last backup will be lost. This must be formally accepted before restore is initiated. |
| RTO | Target 24 hours; actual may be longer if seismic or complex dependencies are in scope. |
| Seismic subproject | The seismic subproject (diskos) and its backing GCS bucket must both be included in the restore scope. Verify this explicitly with Microsoft — the bucket is not automatically included by default. |
| ACL recreation | Seismic subproject ACLs are not automatically restored. The recreation process must be documented and available in the runbook ahead of any restore event. |
| Partition deletion | Accidental data partition deletion is currently not recoverable (soft-delete on roadmap). This is an automatic P1; escalate to senior leadership immediately. |
| Reingestion as alternative | For small-scale deletions, reingestion from source systems by data producer teams is a viable alternative to a full restore. Confirm feasibility with data producer leads before initiating a restore. |
RTO / RPO Summary by Environment
| Environment | RPO | RTO | Notes |
|---|---|---|---|
| Production | ~24 hours (midnight UTC restore) | Target 24 hours | Actual DR test: ~48 hours (extended by seismic issues) |
| Test | N/A for most incidents | Target: same business day | Re-provision from source or re-run test data pipelines |
| Development | N/A | Target: within 1 business day | Team-level resolution; no formal restore mechanism |
RACI Matrix
Role key:
| Code | Role |
|---|---|
| PTL | OSDU Platform Team Lead |
| PT | OSDU Platform Team (engineers) |
| TL | Task Lead |
| MS | Microsoft (ADME Support) |
| SEC | Equinor Security Team |
| DP | Data Producers |
| APP | Application Managers |
RACI key: R = Responsible · A = Accountable · C = Consulted · I = Informed
| Activity | PTL | PT | TL | MS | SEC | DP | APP |
|---|---|---|---|---|---|---|---|
| Monitor alerts / identify anomalous activity | A | R | C | C | |||
| Declare major incident | A/R | C | I | I | I | ||
| Convene War Room / bridge call | R | I | I | I | I | I | |
| Raise Microsoft support ticket (Severity A) | A | R | |||||
| Notify Equinor Security | R | A | |||||
| Instruct pipeline pause (data producers) | A | R | R | I | |||
| Root cause investigation | A | R | I | C | C | C | |
| Decision to DR restore vs. targeted remediation | C | C | A/R | C | |||
| Approve DR restore execution | C | A/R | |||||
| Rotate service principal credentials | A | R | I | C | |||
| Execute DR restore | I | C | I | R | |||
| Platform health checks post-remediation | A | R | I | C | |||
| Data producer validation | I | A | I | R | |||
| End-to-end application validation | I | A | I | C | R | ||
| Incident closure | A | C | R | I | I | ||
| Post-incident report | A | R | I | I | I | ||
| Retrospective / lessons learnt | A/R | C | C | C | C | C | C |
| Runbook / process update | A | R | I |
Open Items — Actions Required Before This Process Is Operational
| # | Action | Owner | Priority |
|---|---|---|---|
| 1 | Implement Azure Monitor alerts for bulk DELETE operations and service health degradation | OSDU Platform Team | High |
| 2 | Set up scheduled Log Analytics record count baseline monitoring | OSDU Platform Team | High |
| 3 | Formally agree impact thresholds and severity levels with the business | OSDU Platform Team + Task Lead | High |
| 4 | Obtain formal RPO acceptance (~24 hours) from the business for production | Task Lead | High |
| 5 | Name a Task Lead deputy for out-of-hours approval authority | Task Lead | High |
| 6 | Populate key contacts table (Section 7.3) and distribute to all stakeholders | OSDU Platform Team Lead | High |
| 7 | Establish on-call rotation for OSDU Platform Team | OSDU Platform Team Lead | High |
| 8 | Document Seismic Store ACL recreation process and add to DR runbook | OSDU Platform Team + Data Producers | High |
| 9 | Run a tabletop exercise against this process with all stakeholders | OSDU Platform Team Lead | Medium |
| 10 | Agree communication cadence and Severity A support ticket process with Microsoft | OSDU Platform Team | Medium |
| 11 | Develop recovery runbook for accidental data partition deletion scenario | OSDU Platform Team + Microsoft | Medium |
| 12 | Define and document on-call escalation paths for out-of-hours P1 incidents | OSDU Platform Team Lead | Medium |
| 13 | Include application consumers in the next DR test to validate end-to-end access | OSDU Platform Team + App Managers | Medium |