Skip to content

DR Trigger & Response Process

ADME / OSDU Platform — Equinor

Document Owner OSDU Platform Team
Version 2.0
Created July 2026
Edited 03 August 2026
Based On DR Test — May 2026 lessons learnt & consolidated results
Frequency of Test Annually, as per requirements set by Equinor
Status Approved- Reviewed by Task Lead

Purpose

This document defines the end-to-end process for detecting, declaring, and executing a Disaster Recovery (DR) restore on the Equinor ADME / OSDU platform. It addresses the key gap identified in the May 2026 DR test: there is currently no defined detection process — no way to identify a mass deletion event or trigger a DR recovery in a real incident.

Scope of applicability: This DR process applies to the production instance only. It does not cover development, test, or staging environments.

Scope — When This Process Applies

This process is intended for platform-level disaster recovery scenarios only, aligned with Equinor's DR requirements. The following disaster scenarios, as defined by Equinor, apply to this platform:

Based on Equinor's requirements for DR testing — see Disaster Scenarios and Mitigation Strategies

Scenario Description
A — Disruptive Cyber-attack A disruptive cyberattack that spreads across Equinor systems and makes them unavailable or encrypted. This includes ransomware attacks, advanced threat actors, and privileged account misuse resulting in intentional deletion of Azure resources and components.
B — Physical Failure A collection of physical or datacenter failures, ranging from failure of a single hardware component to loss of a complete datacenter due to fire or power outage. This includes connectivity issues.
C — Data Deletion Intentional deletion or corruption. Includes privileged account misuse resulting in intentional deletion of data and/or resources such as servers and storage accounts.

Scope — When This Process Does NOT Apply

This process is not a general data correction mechanism. The following scenarios should be handled through normal data operations channels and do not warrant invoking a DR restore:

  • A data producer team accidentally ingested incorrect, duplicate, or malformed records and needs to correct them
  • A bulk ingestion pipeline ran with wrong parameters and data needs to be re-ingested or cleaned up
  • A team soft deleted records they owned as part of a data management activity and wishes to recover them
  • Any error that can be remediated by the data producer re-running their pipeline or re-ingesting from source

Invoking a full DR restore has significant implications — all data written since the last midnight UTC backup will be lost across the entire partition. It should only be triggered when no targeted remediation is possible and the scale of data loss justifies a full rollback.


1. Detection — How to Identify a DR-Triggering Event

1.1 Monitoring Signals to Watch

The following signals should be monitored continuously. Any one of these may indicate a DR-triggering event:

Signal Source What to Look For
Mass record deletion OEP EntitlementsLogs / ADME audit logs Bulk HTTP 204 DELETE responses across one or more OSDU kinds within a short time window
Mass record update ADME Storage API audit logs Large-volume PATCH/PUT operations modifying core fields (ACLs, legal tags, kind) unexpectedly
Record count drop ADME Search API Scheduled query returning record counts significantly below baseline (per OSDU kind)
Seismic subproject deletion Seismic DMS API audit log DELETE on subproject tenant endpoint
Data partition unavailability Azure Monitor / ADME health Partition health check failures or 503 responses from the Storage/Search service
ACL / entitlement group wipe OEP EntitlementsLogs Bulk DELETE of entitlement groups not associated with a planned operation
Unexpected service principal activity Azure AD sign-in logs Unusual or off-hours API activity from known SPNs; signs of compromised credentials
Report from data producer or application manager Direct contact to OSDU Platform Team A data producer or consuming application reports unexpected missing data, failed queries, inaccessible records, or anomalous platform behaviour — treat as high-urgency until ruled out

If you are a data producer or application manager and notice anything unusual — missing records, unexpected query failures, data that should be present but isn't, or any behaviour that suggests data may have been lost or corrupted — contact the OSDU Platform Team immediately and mark your message as high urgency. Do not assume the issue is on your side without first checking with the platform team. Early reporting is critical given the ~24-hour restore window.

1.2 Recommended Alert Setup

Until automated alerting is in place, the following should be implemented as a minimum:

  1. Azure Monitor Alert — set up a metric alert on ADME for a spike in DELETE operations within a rolling 15-minute window.
  2. Log Analytics Query — scheduled query (e.g. every 30 minutes) against OEP EntitlementsLogs and ADME Storage logs, flagging DELETE count > threshold (see Section 2).
  3. Record Count Baseline — maintain a documented baseline count per OSDU kind (reference from May 2026 test: CRS 1,278 / UoM 1,442 / LogCurveType 42,919 / Well 6,304 / Wellbore 12,505). Automated comparison against this baseline should trigger an alert if count drops by more than 10%.
  4. Seismic Subproject Check — periodic API poll to confirm the seismic subproject (diskos) is present and accessible.
  5. Microsoft Fabric Notebook — leverage the existing DR validation notebook (developed during the May 2026 test) as a rapid health check tool that can be run on demand.

Critical constraint: ADME restore runs once per day at midnight UTC. The recovery point is therefore the state of the partition as of the previous midnight. Detection must happen as early as possible to minimise the data loss window within this RPO.

Critical constraint: Any support ticket raised to Microsoft must be classified as Severity A, which mandates round-the-clock support and ensures 24/7 response regardless of time zone.


2. Impact Threshold — When to Declare a DR Incident

Not every deletion event requires a full DR restore. The following thresholds define the trigger levels:

Level Condition
L1 — Monitor < 5% of records deleted in a single kind; appears planned or pipeline-related
L2 — Investigate 5–25% of records in one or more kinds deleted unintentionally; or any seismic subproject deletion
L3 — Declare DR Incident > 25% of records in one or more kinds deleted unintentionally; or full data partition degraded; or seismic subproject + data inaccessible
L4 — Critical Entire data partition unresponsive or deleted; confirmed malicious/accidental wipe

Note: These thresholds are a starting point. They should be formally reviewed and accepted by the business, especially as more data producers and consuming applications go live on the platform.


3. DR Response — Step-by-Step Process

War Room: In the event of a critical outage that requires the DR process to be triggered, all associated roles will be called into a dedicated War Room for fast mobilisation. This ensures all stakeholders are aligned, decisions are made rapidly, and the response is coordinated from a single point of contact.

Phase 1: Detection & Initial Assessment (Target: within 30 minutes)

Step Action Responsible
1.1 Alert fires, team member identifies anomalous deletion / data loss, or a data producer / application manager reports suspicious platform behaviour — all reports treated as high-urgency until investigated Monitoring / Any team member / Data Producers / App Managers
1.2 Run record count checks across all OSDU kinds against baseline OSDU Platform Team
1.3 Check ADME audit logs and OEP EntitlementsLogs to confirm scope and source of deletions OSDU Platform Team
1.4 Determine if deletion is intentional (pipeline run, planned maintenance) or unintentional OSDU Platform Team
1.5 Classify against impact threshold (L1–L4, see Section 2) OSDU Platform Team Lead
1.6 If L3 or L4: Declare DR Incident and notify stakeholders (see Section 4) OSDU Platform Team Lead

Phase 2: Stakeholder Notification & Triage (Target: within 1 hour of declaration)

Step Action Responsible
2.1 Open a dedicated Teams channel or bridge call with all stakeholders OSDU Platform Team Lead
2.2 Notify Microsoft via support ticket — mark as Severity A / Critical for live incidents OSDU Platform Team
2.3 Notify Equinor Security team if compromised credentials are suspected OSDU Platform Team Lead
2.4 Inform data producer teams (Data Producers) — advise them to pause all ingest pipelines to the affected partition OSDU Platform Team
2.5 Notify downstream Application Managers of expected outage window OSDU Platform Team Lead
2.6 Confirm backup state with Microsoft — identify the most recent available restore point (midnight UTC) OSDU Platform Team

Important: All data producer ingest pipelines must be paused before restore is initiated to avoid data written post-incident being overwritten or creating conflicts with the restored state.

Phase 3: Decision to Restore (Target: within 2–3 hours of declaration)

Step Action Responsible
3.1 Confirm restore point with Microsoft — establish what data will be lost (i.e. delta since last midnight UTC backup) OSDU Platform Team
3.2 Obtain formal approval from Task Lead / Data Owner to proceed with restore Task Lead
3.3 Document the decision: scope of restore, data loss accepted, approving stakeholder OSDU Platform Team
3.4 Confirm that Microsoft will initiate restore — request double confirmation step before execution (as practised in May 2026 test) OSDU Platform Team
3.5 If security compromise is suspected: rotate service principal credentials and inform Equinor Security before restore is initiated OSDU Platform Team + Security

Phase 4: Restore Execution (Microsoft-led; target: same day)

Step Action Responsible Time Target
4.1 Microsoft initiates restore to the confirmed restore point Microsoft Per Microsoft SLA
4.2 Maintain hourly call or check-in with Microsoft during active restore window OSDU Platform Team During restore
4.3 Monitor ADME service health for any instability (429s, 500s expected during restore window) OSDU Platform Team During restore
4.4 Do not run any ingest or write operations against the partition during restore All teams During restore

Seismic-specific step: If the seismic subproject (diskos) is in scope, verify with Microsoft that the backing GCS bucket is included in the restore job. This was the root cause of the May 2026 failure — the subproject was restored but the GCS bucket was not. Microsoft have confirmed this has been fixed in the restore job, but it should be explicitly verified.

Phase 5: Validation (Target: within 24 hours of restore completion)

Step Action Responsible
5.1 Run record count checks for all OSDU kinds against baseline OSDU Platform Team
5.2 Verify ACLs and entitlement groups for all restored data kinds OSDU Platform Team
5.3 Verify seismic subproject API accessibility and GCS bucket connectivity OSDU Platform Team + Data Producers
5.4 Run smoke test queries via ADME Search API across all OSDU kinds OSDU Platform Team
5.5 Data producer teams confirm their respective data is accessible and intact Data Producers
5.6 Application Managers confirm end-to-end access Consuming apps / platform team
5.7 Confirm recovery with Task Lead and formally close incident Task Lead + Platform Team Lead

Phase 6: Post-Incident (within 5 business days)

Step Action Responsible
6.1 Produce incident report (timeline, root cause, actions taken, data lost) OSDU Platform Team
6.2 Update runbook with any gaps discovered during the incident OSDU Platform Team
6.3 Conduct retrospective / lessons learnt with all stakeholders OSDU Platform Team Lead
6.4 Raise any new backlog items in ADO / Jira OSDU Platform Team
6.5 Review and update impact thresholds (Section 2) if required OSDU Platform Team + Business

4. Stakeholder Roles & Responsibilities

Role / Team Responsibility
OSDU Platform Team Lead Incident declaration, stakeholder coordination, approval facilitation, Microsoft liaison
OSDU Platform Team (engineers) Detection, log investigation, validation, runbook execution, ACL remediation
Task Lead Formal approval to proceed with restore; business impact assessment
Microsoft (ADME support) Restore execution; GCS bucket restoration (seismic); root cause investigation
Equinor Security Team Credential rotation if compromise suspected; security incident assessment
Data Producer Leads Pause all ingest pipelines to the affected partition on instruction; validate their respective data post-restore
Application Managers Pause non-critical operations during restore; participate in end-to-end validation
App Sec Security assessment and oversight during incident; advise on credential rotation and breach containment
Data Office Data governance oversight; confirm data integrity and compliance requirements post-restore

5. Communication Plan

5.1 Channels & Cadence

Scenario Channel Frequency
Initial declaration Email + Teams message to all stakeholders Once — immediately on declaration
Active restore window Dedicated Teams bridge call or channel Hourly updates while Microsoft is executing
Stakeholder updates (business hours) Teams channel Every 2 hours
Out-of-hours incident On-call contact list (to be defined) As needed
Microsoft communication Severity A support ticket + direct Teams/email Minimum hourly during active restore

5.2 Key Contacts (to be populated)

Role Name Contact
OSDU Platform Team Lead TBD TBD
Task Lead TBD TBD
Microsoft Account Contact TBD TBD
Microsoft Support (Severity A) Azure Portal — support ticket aka.ms/azuresupport
Equinor Security On-Call TBD TBD
Data Producer Leads TBD TBD
Application Managers TBD TBD
Data Office TBD TBD
App Sec TBD TBD

6. RTO / RPO Reference

Metric Current State Notes
RPO (Recovery Point Objective) ~24 hours Restore runs once daily at midnight UTC — data written after the last backup is lost
RTO (Recovery Time Objective) Target: 24 hours Actual in May 2026 test: ~48 hours (extended by Seismic subproject issues)
Time to detect Target: < 30 minutes Dependent on alert setup (see Section 1.2)
Time to declare Target: < 1 hour from detection Requires Task Lead availability — a deputy should be named
Time to begin restore Target: 2–3 hours from declaration Microsoft must be engaged quickly; USA time zone offset may delay execution
Time to complete restore ~2–3 hours (proven in test) Microsoft-executed; exclude Seismic subproject complications
Time to validate Target: < 24 hours post-restore Extended in test due to Seismic bucket issue — mitigated going forward

RPO acceptance required: The ~24-hour RPO must be formally accepted by the business. If a tighter RPO is required, this should be raised with Microsoft as a platform configuration change request.


7. Special Considerations

7.1 Seismic Store (Data Producers / DISKOS)

  • The seismic subproject (diskos) requires special handling — both the subproject metadata and the backing GCS bucket must be restored.
  • Verify explicitly with Microsoft that the GCS bucket is included in the restore scope.
  • ACL entitlement groups for the seismic subproject are not automatically restored — manual recreation via CLI may be required (process to be documented in runbook by OSDU Platform Team + Data Producers).

7.2 Security Response

  • If the deletion event is suspected to be the result of a compromised service principal, credentials should be rotated before restore is initiated.
  • The Equinor Security team should be looped in to assess the scope of any potential breach.
  • Reference: OEP EntitlementsLogs and Azure AD sign-in logs are the primary sources for investigating suspicious SPN activity.

7.3 Data Partition Deletion

  • Microsoft have confirmed that accidental data partition deletion is not currently recoverable (soft-delete is on the roadmap).
  • If the data partition itself is deleted, this is an L4 critical incident and senior leadership must be engaged immediately.
  • This scenario requires a separate recovery runbook — to be developed with Microsoft.

7.4 Time Zone Considerations

  • Microsoft restore execution is USA-based. Requests raised late in the European day may not be actioned until the following business day.
  • For live incidents, use Severity A / Critical support ticket classification to ensure 24/7 response.
  • Target to raise the Microsoft support ticket before midday CET where possible to avoid overnight delays.

7.5 Reingestion as an Alternative

  • For small-scale deletions (L1–L2) where a full restore is disproportionate, reingestion of deleted records from source systems (Data Producers) is a viable alternative.
  • This was validated in the May 2026 test (10 LogCurveType records reingested successfully).
  • Data producer teams should confirm whether their source pipelines can be re-run selectively.

8. Open Items — Actions Required Before This Process Is Operational

# Action Owner Priority
1 Implement Azure Monitor alert for bulk DELETE operations on ADME OSDU Platform Team High
2 Set up scheduled Log Analytics query for record count baseline monitoring OSDU Platform Team High
3 Formally agree and document impact thresholds (Section 2) with the business OSDU Platform Team + Task Lead High
4 Obtain formal RPO acceptance (~24 hours) from the business Task Lead High
5 Name a Task Lead deputy for out-of-hours approval authority Task Lead High
6 Populate key contacts table (Section 5.3) OSDU Platform Team Lead High
7 Establish on-call rotation for OSDU Platform Team OSDU Platform Team Lead High
8 Document Seismic Store ACL recreation process (with Data Producers) and add to runbook OSDU Platform Team + Data Producers High
9 Agree communication cadence and support ticket process with Microsoft for live DR events OSDU Platform Team Medium
10 Develop recovery runbook for accidental data partition deletion scenario OSDU Platform Team + Microsoft Medium
11 Run a tabletop exercise against this process before the next DR test OSDU Platform Team Medium

9. RACI Matrix

Role key:

Code Role
PTL OSDU Platform Team Lead
PT OSDU Platform Team (engineers)
PO Task Lead
MS Microsoft (ADME Support)
SEC Equinor Security Team
DP Data Producers
APP Application Managers

RACI key: R = Responsible · A = Accountable · C = Consulted · I = Informed


Phase 1 — Detection & Initial Assessment

Activity PTL PT PO MS SEC Data Producers APP
Monitor alerts / identify anomalous deletion A R
Run record count checks against baseline A R
Review ADME audit logs & EntitlementsLogs A R
Classify incident level (L1–L4) R/A C I
Declare DR Incident R/A I I I I

Phase 2 — Stakeholder Notification & Triage

Activity PTL PT PO MS SEC Data Producers APP
Open Teams bridge / incident channel R/A C I I I
Raise Microsoft Severity A support ticket A R I
Notify Security team (if compromise suspected) R/A I I
Instruct data producers to pause ingest pipelines A R I I
Notify Application Managers of outage window R/A I I
Confirm backup state & restore point with Microsoft A R I C

Phase 3 — Decision to Restore

Activity PTL PT PO MS SEC Data Producers APP
Confirm scope and data loss from restore point delta A R C C
Obtain formal approval to proceed with restore C R/A
Document approval decision (scope, data loss, approver) A R I
Rotate SPN credentials (if compromise suspected) A R I C
Request double-confirmation from Microsoft before execution R/A C I I

Phase 4 — Restore Execution

Activity PTL PT PO MS SEC Data Producers APP
Initiate and execute restore C C I R/A
Verify seismic GCS bucket is included in restore scope A R C C
Monitor ADME service health during restore window A R C
Maintain hourly check-in with Microsoft during restore R/A C I R
Enforce write/ingest freeze on partition A R I I

Phase 5 — Validation

Activity PTL PT PO MS SEC Data Producers APP
Record count checks (all OSDU kinds vs. baseline) A R
ACL / entitlement group verification A R
Seismic subproject API & GCS bucket validation A R C R
Seismic ACL recreation (if required) A R C
Data Producers data validation C I R/A
End-to-end application manager validation I C A R
Formal incident closure sign-off R I A I I

Phase 6 — Post-Incident

Activity PTL PT PO MS SEC Data Producers APP
Produce incident report (timeline, root cause, data loss) A R I C C I I
Update DR runbook with gaps discovered A R C C
Update Seismic Store runbook section A R C
Conduct lessons learnt retrospective R/A C C C C C
Raise backlog items in ADO / Jira A R C
Review and update impact thresholds if required A C R C

This document should be reviewed and updated following each DR test or live incident. Version history to be maintained alongside the DR runbook.


Last update: 2026-08-17