Skip to content

Major Incident Response Process

ADME / OSDU Platform — Equinor

Document Owner OSDU Platform Team
Version 1.0
Created August 2026
Based On DR Test May 2026 — Lessons Learnt & DR Trigger & Response Process
Status Draft

Purpose

This document defines the end-to-end process for detecting, declaring, and responding to a major incident on the Equinor ADME / OSDU platform. It applies across all environments — development, test, and production — and is intended to provide a consistent, repeatable response framework regardless of the specific incident type.

A major incident is any event that causes, or has the potential to cause, significant disruption to platform availability, data integrity, or the ability of data producers and consuming applications to operate. A Disaster Recovery (DR) restore is one example of a major incident; other examples include large-scale data corruption, platform-wide service outages, and confirmed security breaches.

This document does not replace the DR-specific runbook. For DR restore steps on the production instance, refer to the DR Trigger & Response Process document.

Scope

Environments Covered

Environment Examples Incident Handling
Production (data) Live platform serving all data producers and consuming applications Full incident process applies; highest urgency
Test (test, staging) Pre-production validation, DR rehearsals, integration testing Full incident process applies where data integrity or test fidelity is at risk; lower urgency than production
Development Developer sandboxes, feature branches, spike environments Lightweight process applies; team-level resolution preferred before escalation

Incident Types Covered

The following categories constitute a major incident on this platform:

Category Examples
Data Loss / Corruption Malicious or confirmed unauthorised deletion of records, deletion of a data partition, deliberate corruption of ACLs or legal tags
Platform Availability Full or partial unavailability of ADME services (Storage, Search, Indexer, Seismic DMS, Entitlements), data partition unresponsive for more than 30 minutes to 1 hour
Security Breach Compromised service principal credentials, unauthorised access, suspected ransomware or cyberattack, privilege escalation
Infrastructure Failure Azure region/datacenter failure, backing storage (GCS/blob) inaccessible, network connectivity loss to ADME
Dependency Failure Critical upstream or downstream service failure affecting data flow or consuming applications (e.g. SMDA, SDMC, identity provider)

What This Process Does NOT Cover

  • Data corrections that can be addressed by a data producer re-running their pipeline
  • Planned maintenance or environment refreshes
  • Individual pipeline failures or ingestion errors with no broader platform impact
  • Feature bugs or configuration changes in a single service that do not affect data integrity or platform availability

Incident Severity Levels

Level Name Definition Response Target
P1 — Critical Production environment severely degraded, significant data loss or corruption confirmed, or active security breach Immediate response; War Room convened within 30 minutes; 24/7 until resolved
P2 — Major Production environment partially degraded, or data loss contained to a subset of kinds/partitions, or test environment fully unavailable ahead of a scheduled test Response within 1 hour; dedicated coordination channel opened
P3 — Significant Test or dev environment degraded; production impact contained or under investigation; potential data integrity concern Response within 4 hours during business hours
P4 — Minor Dev environment issue; no production or test impact; investigation required but no immediate risk Response within 1 business day

Environment modifier: Any incident directly impacting the production data partition is automatically elevated to P1 or P2 regardless of apparent scope, until confirmed otherwise. If a Microsoft support ticket is required, it must be raised at the same severity level (Severity A / Critical for P1; Severity B for P2) and must not be downgraded without explicit agreement from the OSDU Platform Team Lead.


Detection — How to Identify a Major Incident

Monitoring Signals

The following signals may indicate a major incident. Any team member or data producer who observes these should raise immediately with the OSDU Platform Team:

Signal Source Potential Incident Type
Mass record deletion or unexpected record count drop ADME Search API / audit logs Data Loss / DR Event
Bulk ACL or entitlement group deletion OEP EntitlementsLogs Data Loss / Security
Service returning repeated 500 / 503 errors ADME health monitoring / Azure Monitor Platform Availability
Seismic subproject inaccessible or missing Seismic DMS API Data Loss / Platform Availability
Data partition unresponsive or deleted Azure Portal / ADME health DR Event / Infrastructure Failure
Unusual or off-hours service principal activity Azure AD sign-in logs Security Breach
Downstream application reporting unexpected missing or inaccessible data Direct report from application or data producer Data Loss / Platform Availability
Pipeline failures across multiple data producer teams simultaneously Data producer monitoring Platform Availability / Infrastructure Failure
Azure region or datacenter health alert Azure Service Health Infrastructure Failure

Reporting by Data Producers and Application Managers

⚠️ If you are a data producer or application manager and notice anything unusual — missing records, unexpected query failures, data that should be present but isn't, or behaviour that suggests data may have been lost or corrupted — contact the OSDU Platform Team immediately and mark your message as high urgency. Do not assume the issue is on your side without first checking with the platform team. Early detection is critical.

Recommended Alert Setup

The following should be in place as a minimum on the production instance:

  1. Azure Monitor Alert — metric alert for a spike in HTTP 4xx/5xx error rates or DELETE operations within a rolling 15-minute window.
  2. Log Analytics Scheduled Query — runs every 30 minutes against ADME Storage and OEP EntitlementsLogs, flagging bulk DELETEs above a defined threshold.
  3. Record Count Baseline Check — automated comparison of record counts per OSDU kind against a documented baseline; triggers alert on >10% drop.
  4. Azure Service Health Alerts — configured to notify the team of Azure region or ADME service health events.
  5. Seismic Subproject Poll — periodic check confirming subproject presence and GCS bucket accessibility.

Incident Declaration

Who Can Declare

Any OSDU Platform Team member can initiate an incident investigation. Only the OSDU Platform Team Lead (or their named deputy) can formally declare a major incident.

Declaration Criteria

Declare a major incident if any of the following are true:

  • Data loss or corruption is confirmed or strongly suspected at scale (see Section 2 for severity)
  • A production ADME service is unavailable and the cause is not immediately resolvable
  • A security breach or compromise is suspected
  • A test environment is non-functional and a scheduled test or milestone is imminent
  • Equinor's DR disaster scenarios (Cyber-attack / Physical Failure / Data Deletion) are triggered

When in doubt, declare. It is always better to stand down from an incident than to discover one was not declared in time.


Response Phases

📢 War Room: For P1/P2 incidents, all associated roles will be called into a dedicated War Room (Teams bridge call or channel) immediately on declaration. This ensures all stakeholders are aligned, decisions are made rapidly, and the response is coordinated from a single point.

Detection & Initial Assessment (Target: within 30 minutes)

Step Action Responsible Time Target
1.1 Alert fires, anomaly observed, or report received from data producer / application manager — treat all reports as high-urgency until investigated Monitoring system / Any team member T+0
1.2 Assign an incident lead to coordinate the initial response OSDU Platform Team Lead T+0 to T+10 min
1.3 Run initial health checks: record counts, service status, audit logs, Azure Portal OSDU Platform Team T+0 to T+15 min
1.4 Determine environment affected (prod / test / dev) and confirm whether incident is real or a false positive OSDU Platform Team T+15 min
1.5 Classify against severity level (P1–P4) and incident type (Section 2) OSDU Platform Team Lead T+15 to T+30 min
1.6 Formally declare major incident (if applicable) and open coordination channel OSDU Platform Team Lead T+30 min

Phase 2: Stakeholder Notification & Triage (Target: within 1 hour of declaration)

Step Action Responsible Time Target
2.1 Open dedicated Teams bridge call / channel; invite all stakeholders per Section 6 OSDU Platform Team Lead T+30 min
2.2 Send initial incident notification (environment, nature of incident, current status) OSDU Platform Team Lead T+30 to T+45 min
2.3 Engage Microsoft via support ticket if ADME platform action is required — classify as Severity A / Critical for P1/P2 OSDU Platform Team T+30 to T+45 min
2.4 Notify Equinor Security Team if a breach or credential compromise is suspected OSDU Platform Team Lead T+30 to T+45 min
2.5 Instruct data producer teams to pause all ingest pipelines to the affected partition — P1/P2 only OSDU Platform Team T+45 min to T+1 hr
2.6 Notify downstream Application Managers of expected impact and outage window OSDU Platform Team Lead T+1 hr

Phase 3: Investigation & Decision (Target: within 2–3 hours of declaration)

Step Action Responsible Time Target
3.1 Investigate root cause — review ADME audit logs, Azure AD sign-in logs, OEP EntitlementsLogs, and Azure Monitor OSDU Platform Team T+1 to T+2 hr
3.2 Determine whether the incident can be resolved through targeted remediation (reingestion, manual fix, config rollback) or requires a full DR restore OSDU Platform Team Lead T+2 hr
3.3 If DR restore is required: follow the DR Trigger & Response Process; obtain Task Lead approval before proceeding Task Lead + Platform Team Lead T+2 hr
3.4 If security breach is suspected: rotate service principal credentials and confirm with Equinor Security before any restore or repair action Platform Team + Equinor Security T+2 hr
3.5 Document the decision: scope, affected data, chosen remediation approach, approving stakeholder OSDU Platform Team T+2 hr
3.6 Confirm remediation plan with Microsoft if their involvement is required OSDU Platform Team T+2 to T+3 hr

Phase 4: Remediation Execution

Step Action Responsible Time Target
4.1 Execute agreed remediation — DR restore, service restart, config rollback, data reingestion, or security remediation OSDU Platform Team / Microsoft Per agreed plan
4.2 Provide regular updates (minimum hourly) to all stakeholders in the coordination channel OSDU Platform Team Lead During execution
4.3 Do not run write or ingest operations against the affected partition during active remediation All teams During execution
4.4 Monitor for unexpected side effects — watch for 429s, 500s, or secondary failures OSDU Platform Team During execution

Phase 5: Validation (Target: within 24 hours of remediation completion)

Step Action Responsible Time Target
5.1 Run platform health checks: service availability, record count checks against baseline, ACL verification OSDU Platform Team T+0 to T+2 hr post-remediation
5.2 Run smoke test queries across affected OSDU data kinds OSDU Platform Team T+2 to T+4 hr post-remediation
5.3 Data producer teams confirm their respective data is accessible and intact Data Producers T+4 to T+8 hr post-remediation
5.4 Application Managers confirm end-to-end application access Application Managers T+8 to T+24 hr post-remediation
5.5 If DR restore was executed: verify seismic subproject and GCS bucket accessibility explicitly OSDU Platform Team + SDMC T+2 hr post-restore
5.6 Confirm recovery with Task Lead and formally close the incident Task Lead + Platform Team Lead T+24 hr post-remediation

Phase 6: Post-Incident (within 5 business days)

Step Action Responsible
6.1 Produce incident report: timeline, root cause, scope of impact, actions taken, data lost (if any) OSDU Platform Team
6.2 Conduct retrospective / lessons learnt session with all stakeholders OSDU Platform Team Lead
6.3 Update runbooks and this process document with gaps discovered OSDU Platform Team
6.4 Raise backlog items in ADO / Jira for follow-up actions OSDU Platform Team
6.5 Review and update severity thresholds and escalation paths if required OSDU Platform Team + Business

Roles & Responsibilities

Role Responsibility Notification Priority
OSDU Platform Team Lead Incident declaration, stakeholder coordination, Microsoft liaison, approval facilitation Immediate — T+0
OSDU Platform Team (engineers) Detection, log investigation, runbook execution, health checks, ACL remediation Immediate — T+0
Task Lead Formal approval for DR restore or major remediation actions; business impact assessment; incident closure Within 1 hour of P1/P2 declaration
Microsoft (ADME Support) Restore execution (DR events); root cause investigation; GCS bucket and infrastructure remediation Within 30–45 min via Severity A ticket (P1/P2)
Equinor Security Team Credential rotation; security incident assessment; breach containment Within 1 hour if compromise suspected
Data Producer Leads Pause ingest pipelines on instruction; validate respective data post-remediation Within 1 hour of declaration
Application Managers Pause non-critical operations during remediation; participate in end-to-end validation Notified within 1 hour; active from Phase 5
App Sec Security assessment and oversight; advise on credential rotation and breach scope Within 1 hour if compromise suspected
Data Office Data governance oversight; confirm data integrity and compliance post-incident Notified within 1 hour of P1/P2 declaration

Communication Plan

Channels & Cadence

Scenario Channel Frequency
Initial declaration Email + Teams message to all stakeholders Once — immediately on declaration
Active remediation Dedicated Teams bridge call or channel Hourly updates while remediation is in progress
Stakeholder updates (business hours) Teams channel Every 2 hours
Out-of-hours P1 incident On-call contact list (to be defined) As needed; continuous until resolved
Microsoft communication Severity A support ticket + direct Teams/email Minimum hourly during active engagement

Communication Templates

Initial Declaration Message:

🔴 [MAJOR INCIDENT DECLARED — P{severity}] Environment: {prod / test / dev} Nature: {brief description of incident type and impact} Status: Under investigation Incident Lead: {name} Bridge / Channel: {link} Next update: {time}

Incident Closure Message:

✅ [INCIDENT CLOSED — P{severity}] Resolution: {brief description of remediation applied} Impact summary: {data affected, downtime, etc.} Incident report: {link — to follow within 5 business days} Retrospective: {date/time — to be scheduled}

Key Contacts (to be populated)

Role Name Contact
OSDU Platform Team Lead TBD TBD
Task Lead TBD TBD
Task Lead Deputy (out-of-hours) TBD TBD
Microsoft Account Contact TBD TBD
Microsoft Support (Severity A) Azure Portal — support ticket aka.ms/azuresupport
Equinor Security On-Call TBD TBD
Data Producer Leads TBD TBD
Application Managers TBD TBD
Data Office TBD TBD
App Sec TBD TBD

Environment-Specific Considerations

Production (data partition)

  • Full incident process applies without exception.
  • Any P1/P2 incident must involve Task Lead approval before major remediation actions are taken.
  • Attempt targeted remediation before escalating to more disruptive interventions (e.g. full restores, partition-level actions).
  • Where Microsoft support is required, raise the ticket before midday CET where possible to avoid overnight delays due to USA-based support. For P1 incidents, Severity A tickets guarantee 24/7 response.

Test Environments (test, staging)

  • Full incident process applies where data integrity or test fidelity is at risk (e.g. ahead of a scheduled DR rehearsal or certification activity).
  • Test environments are isolated from production — partition isolation must be verified before concluding there is no production impact.
  • A simpler, team-level resolution path is acceptable for low-risk test environment issues (P3/P4).
  • For planned DR tests, a separate test runbook governs the execution; this process governs any unplanned failures during or after the test.

Development Environments

  • Team-level resolution is the default. Escalate to OSDU Platform Team only if the issue cannot be resolved within the team.
  • Development environments must not be used as a workaround for production issues.

DR-Specific Considerations

When a major incident results in a DR restore being required, the following apply in addition to this general process:

Topic Consideration
RPO ~24 hours — restore runs once daily at midnight UTC. All data written since the last backup will be lost. This must be formally accepted before restore is initiated.
RTO Target 24 hours; actual may be longer if seismic or complex dependencies are in scope.
Seismic subproject The seismic subproject (diskos) and its backing GCS bucket must both be included in the restore scope. Verify this explicitly with Microsoft — the bucket is not automatically included by default.
ACL recreation Seismic subproject ACLs are not automatically restored. The recreation process must be documented and available in the runbook ahead of any restore event.
Partition deletion Accidental data partition deletion is currently not recoverable (soft-delete on roadmap). This is an automatic P1; escalate to senior leadership immediately.
Reingestion as alternative For small-scale deletions, reingestion from source systems by data producer teams is a viable alternative to a full restore. Confirm feasibility with data producer leads before initiating a restore.

RTO / RPO Summary by Environment

Environment RPO RTO Notes
Production ~24 hours (midnight UTC restore) Target 24 hours Actual DR test: ~48 hours (extended by seismic issues)
Test N/A for most incidents Target: same business day Re-provision from source or re-run test data pipelines
Development N/A Target: within 1 business day Team-level resolution; no formal restore mechanism

RACI Matrix

Role key:

Code Role
PTL OSDU Platform Team Lead
PT OSDU Platform Team (engineers)
TL Task Lead
MS Microsoft (ADME Support)
SEC Equinor Security Team
DP Data Producers
APP Application Managers

RACI key: R = Responsible · A = Accountable · C = Consulted · I = Informed

Activity PTL PT TL MS SEC DP APP
Monitor alerts / identify anomalous activity A R C C
Declare major incident A/R C I I I
Convene War Room / bridge call R I I I I I
Raise Microsoft support ticket (Severity A) A R
Notify Equinor Security R A
Instruct pipeline pause (data producers) A R R I
Root cause investigation A R I C C C
Decision to DR restore vs. targeted remediation C C A/R C
Approve DR restore execution C A/R
Rotate service principal credentials A R I C
Execute DR restore I C I R
Platform health checks post-remediation A R I C
Data producer validation I A I R
End-to-end application validation I A I C R
Incident closure A C R I I
Post-incident report A R I I I
Retrospective / lessons learnt A/R C C C C C C
Runbook / process update A R I

Open Items — Actions Required Before This Process Is Operational

# Action Owner Priority
1 Implement Azure Monitor alerts for bulk DELETE operations and service health degradation OSDU Platform Team High
2 Set up scheduled Log Analytics record count baseline monitoring OSDU Platform Team High
3 Formally agree impact thresholds and severity levels with the business OSDU Platform Team + Task Lead High
4 Obtain formal RPO acceptance (~24 hours) from the business for production Task Lead High
5 Name a Task Lead deputy for out-of-hours approval authority Task Lead High
6 Populate key contacts table (Section 7.3) and distribute to all stakeholders OSDU Platform Team Lead High
7 Establish on-call rotation for OSDU Platform Team OSDU Platform Team Lead High
8 Document Seismic Store ACL recreation process and add to DR runbook OSDU Platform Team + Data Producers High
9 Run a tabletop exercise against this process with all stakeholders OSDU Platform Team Lead Medium
10 Agree communication cadence and Severity A support ticket process with Microsoft OSDU Platform Team Medium
11 Develop recovery runbook for accidental data partition deletion scenario OSDU Platform Team + Microsoft Medium
12 Define and document on-call escalation paths for out-of-hours P1 incidents OSDU Platform Team Lead Medium
13 Include application consumers in the next DR test to validate end-to-end access OSDU Platform Team + App Managers Medium

Last update: 2026-09-17