DR Trigger & Response Process
ADME / OSDU Platform — Equinor
|
|
| Document Owner |
OSDU Platform Team |
| Version |
2.0 |
| Created |
July 2026 |
| Edited |
03 August 2026 |
| Based On |
DR Test — May 2026 lessons learnt & consolidated results |
| Frequency of Test |
Annually, as per requirements set by Equinor |
| Status |
Approved- Reviewed by Task Lead |
Purpose
This document defines the end-to-end process for detecting, declaring, and executing a Disaster Recovery (DR) restore on the Equinor ADME / OSDU platform. It addresses the key gap identified in the May 2026 DR test: there is currently no defined detection process — no way to identify a mass deletion event or trigger a DR recovery in a real incident.
Scope of applicability: This DR process applies to the production instance only. It does not cover development, test, or staging environments.
Scope — When This Process Applies
This process is intended for platform-level disaster recovery scenarios only, aligned with Equinor's DR requirements. The following disaster scenarios, as defined by Equinor, apply to this platform:
Based on Equinor's requirements for DR testing — see Disaster Scenarios and Mitigation Strategies
| Scenario |
Description |
| A — Disruptive Cyber-attack |
A disruptive cyberattack that spreads across Equinor systems and makes them unavailable or encrypted. This includes ransomware attacks, advanced threat actors, and privileged account misuse resulting in intentional deletion of Azure resources and components. |
| B — Physical Failure |
A collection of physical or datacenter failures, ranging from failure of a single hardware component to loss of a complete datacenter due to fire or power outage. This includes connectivity issues. |
| C — Data Deletion |
Intentional deletion or corruption. Includes privileged account misuse resulting in intentional deletion of data and/or resources such as servers and storage accounts. |
Scope — When This Process Does NOT Apply
This process is not a general data correction mechanism. The following scenarios should be handled through normal data operations channels and do not warrant invoking a DR restore:
- A data producer team accidentally ingested incorrect, duplicate, or malformed records and needs to correct them
- A bulk ingestion pipeline ran with wrong parameters and data needs to be re-ingested or cleaned up
- A team soft deleted records they owned as part of a data management activity and wishes to recover them
- Any error that can be remediated by the data producer re-running their pipeline or re-ingesting from source
Invoking a full DR restore has significant implications — all data written since the last midnight UTC backup will be lost across the entire partition. It should only be triggered when no targeted remediation is possible and the scale of data loss justifies a full rollback.
1. Detection — How to Identify a DR-Triggering Event
1.1 Monitoring Signals to Watch
The following signals should be monitored continuously. Any one of these may indicate a DR-triggering event:
| Signal |
Source |
What to Look For |
| Mass record deletion |
OEP EntitlementsLogs / ADME audit logs |
Bulk HTTP 204 DELETE responses across one or more OSDU kinds within a short time window |
| Mass record update |
ADME Storage API audit logs |
Large-volume PATCH/PUT operations modifying core fields (ACLs, legal tags, kind) unexpectedly |
| Record count drop |
ADME Search API |
Scheduled query returning record counts significantly below baseline (per OSDU kind) |
| Seismic subproject deletion |
Seismic DMS API audit log |
DELETE on subproject tenant endpoint |
| Data partition unavailability |
Azure Monitor / ADME health |
Partition health check failures or 503 responses from the Storage/Search service |
| ACL / entitlement group wipe |
OEP EntitlementsLogs |
Bulk DELETE of entitlement groups not associated with a planned operation |
| Unexpected service principal activity |
Azure AD sign-in logs |
Unusual or off-hours API activity from known SPNs; signs of compromised credentials |
| Report from data producer or application manager |
Direct contact to OSDU Platform Team |
A data producer or consuming application reports unexpected missing data, failed queries, inaccessible records, or anomalous platform behaviour — treat as high-urgency until ruled out |
If you are a data producer or application manager and notice anything unusual — missing records, unexpected query failures, data that should be present but isn't, or any behaviour that suggests data may have been lost or corrupted — contact the OSDU Platform Team immediately and mark your message as high urgency. Do not assume the issue is on your side without first checking with the platform team. Early reporting is critical given the ~24-hour restore window.
1.2 Recommended Alert Setup
Until automated alerting is in place, the following should be implemented as a minimum:
- Azure Monitor Alert — set up a metric alert on ADME for a spike in DELETE operations within a rolling 15-minute window.
- Log Analytics Query — scheduled query (e.g. every 30 minutes) against
OEP EntitlementsLogs and ADME Storage logs, flagging DELETE count > threshold (see Section 2).
- Record Count Baseline — maintain a documented baseline count per OSDU kind (reference from May 2026 test: CRS 1,278 / UoM 1,442 / LogCurveType 42,919 / Well 6,304 / Wellbore 12,505). Automated comparison against this baseline should trigger an alert if count drops by more than 10%.
- Seismic Subproject Check — periodic API poll to confirm the seismic subproject (
diskos) is present and accessible.
- Microsoft Fabric Notebook — leverage the existing DR validation notebook (developed during the May 2026 test) as a rapid health check tool that can be run on demand.
Critical constraint: ADME restore runs once per day at midnight UTC. The recovery point is therefore the state of the partition as of the previous midnight. Detection must happen as early as possible to minimise the data loss window within this RPO.
Critical constraint: Any support ticket raised to Microsoft must be classified as Severity A, which mandates round-the-clock support and ensures 24/7 response regardless of time zone.
2. Impact Threshold — When to Declare a DR Incident
Not every deletion event requires a full DR restore. The following thresholds define the trigger levels:
| Level |
Condition |
| L1 — Monitor |
< 5% of records deleted in a single kind; appears planned or pipeline-related |
| L2 — Investigate |
5–25% of records in one or more kinds deleted unintentionally; or any seismic subproject deletion |
| L3 — Declare DR Incident |
> 25% of records in one or more kinds deleted unintentionally; or full data partition degraded; or seismic subproject + data inaccessible |
| L4 — Critical |
Entire data partition unresponsive or deleted; confirmed malicious/accidental wipe |
Note: These thresholds are a starting point. They should be formally reviewed and accepted by the business, especially as more data producers and consuming applications go live on the platform.
3. DR Response — Step-by-Step Process
War Room: In the event of a critical outage that requires the DR process to be triggered, all associated roles will be called into a dedicated War Room for fast mobilisation. This ensures all stakeholders are aligned, decisions are made rapidly, and the response is coordinated from a single point of contact.
Phase 1: Detection & Initial Assessment (Target: within 30 minutes)
| Step |
Action |
Responsible |
| 1.1 |
Alert fires, team member identifies anomalous deletion / data loss, or a data producer / application manager reports suspicious platform behaviour — all reports treated as high-urgency until investigated |
Monitoring / Any team member / Data Producers / App Managers |
| 1.2 |
Run record count checks across all OSDU kinds against baseline |
OSDU Platform Team |
| 1.3 |
Check ADME audit logs and OEP EntitlementsLogs to confirm scope and source of deletions |
OSDU Platform Team |
| 1.4 |
Determine if deletion is intentional (pipeline run, planned maintenance) or unintentional |
OSDU Platform Team |
| 1.5 |
Classify against impact threshold (L1–L4, see Section 2) |
OSDU Platform Team Lead |
| 1.6 |
If L3 or L4: Declare DR Incident and notify stakeholders (see Section 4) |
OSDU Platform Team Lead |
Phase 2: Stakeholder Notification & Triage (Target: within 1 hour of declaration)
| Step |
Action |
Responsible |
| 2.1 |
Open a dedicated Teams channel or bridge call with all stakeholders |
OSDU Platform Team Lead |
| 2.2 |
Notify Microsoft via support ticket — mark as Severity A / Critical for live incidents |
OSDU Platform Team |
| 2.3 |
Notify Equinor Security team if compromised credentials are suspected |
OSDU Platform Team Lead |
| 2.4 |
Inform data producer teams (Data Producers) — advise them to pause all ingest pipelines to the affected partition |
OSDU Platform Team |
| 2.5 |
Notify downstream Application Managers of expected outage window |
OSDU Platform Team Lead |
| 2.6 |
Confirm backup state with Microsoft — identify the most recent available restore point (midnight UTC) |
OSDU Platform Team |
Important: All data producer ingest pipelines must be paused before restore is initiated to avoid data written post-incident being overwritten or creating conflicts with the restored state.
Phase 3: Decision to Restore (Target: within 2–3 hours of declaration)
| Step |
Action |
Responsible |
| 3.1 |
Confirm restore point with Microsoft — establish what data will be lost (i.e. delta since last midnight UTC backup) |
OSDU Platform Team |
| 3.2 |
Obtain formal approval from Task Lead / Data Owner to proceed with restore |
Task Lead |
| 3.3 |
Document the decision: scope of restore, data loss accepted, approving stakeholder |
OSDU Platform Team |
| 3.4 |
Confirm that Microsoft will initiate restore — request double confirmation step before execution (as practised in May 2026 test) |
OSDU Platform Team |
| 3.5 |
If security compromise is suspected: rotate service principal credentials and inform Equinor Security before restore is initiated |
OSDU Platform Team + Security |
Phase 4: Restore Execution (Microsoft-led; target: same day)
| Step |
Action |
Responsible |
Time Target |
| 4.1 |
Microsoft initiates restore to the confirmed restore point |
Microsoft |
Per Microsoft SLA |
| 4.2 |
Maintain hourly call or check-in with Microsoft during active restore window |
OSDU Platform Team |
During restore |
| 4.3 |
Monitor ADME service health for any instability (429s, 500s expected during restore window) |
OSDU Platform Team |
During restore |
| 4.4 |
Do not run any ingest or write operations against the partition during restore |
All teams |
During restore |
Seismic-specific step: If the seismic subproject (diskos) is in scope, verify with Microsoft that the backing GCS bucket is included in the restore job. This was the root cause of the May 2026 failure — the subproject was restored but the GCS bucket was not. Microsoft have confirmed this has been fixed in the restore job, but it should be explicitly verified.
Phase 5: Validation (Target: within 24 hours of restore completion)
| Step |
Action |
Responsible |
| 5.1 |
Run record count checks for all OSDU kinds against baseline |
OSDU Platform Team |
| 5.2 |
Verify ACLs and entitlement groups for all restored data kinds |
OSDU Platform Team |
| 5.3 |
Verify seismic subproject API accessibility and GCS bucket connectivity |
OSDU Platform Team + Data Producers |
| 5.4 |
Run smoke test queries via ADME Search API across all OSDU kinds |
OSDU Platform Team |
| 5.5 |
Data producer teams confirm their respective data is accessible and intact |
Data Producers |
| 5.6 |
Application Managers confirm end-to-end access |
Consuming apps / platform team |
| 5.7 |
Confirm recovery with Task Lead and formally close incident |
Task Lead + Platform Team Lead |
Phase 6: Post-Incident (within 5 business days)
| Step |
Action |
Responsible |
| 6.1 |
Produce incident report (timeline, root cause, actions taken, data lost) |
OSDU Platform Team |
| 6.2 |
Update runbook with any gaps discovered during the incident |
OSDU Platform Team |
| 6.3 |
Conduct retrospective / lessons learnt with all stakeholders |
OSDU Platform Team Lead |
| 6.4 |
Raise any new backlog items in ADO / Jira |
OSDU Platform Team |
| 6.5 |
Review and update impact thresholds (Section 2) if required |
OSDU Platform Team + Business |
4. Stakeholder Roles & Responsibilities
| Role / Team |
Responsibility |
| OSDU Platform Team Lead |
Incident declaration, stakeholder coordination, approval facilitation, Microsoft liaison |
| OSDU Platform Team (engineers) |
Detection, log investigation, validation, runbook execution, ACL remediation |
| Task Lead |
Formal approval to proceed with restore; business impact assessment |
| Microsoft (ADME support) |
Restore execution; GCS bucket restoration (seismic); root cause investigation |
| Equinor Security Team |
Credential rotation if compromise suspected; security incident assessment |
| Data Producer Leads |
Pause all ingest pipelines to the affected partition on instruction; validate their respective data post-restore |
| Application Managers |
Pause non-critical operations during restore; participate in end-to-end validation |
| App Sec |
Security assessment and oversight during incident; advise on credential rotation and breach containment |
| Data Office |
Data governance oversight; confirm data integrity and compliance requirements post-restore |
5. Communication Plan
5.1 Channels & Cadence
| Scenario |
Channel |
Frequency |
| Initial declaration |
Email + Teams message to all stakeholders |
Once — immediately on declaration |
| Active restore window |
Dedicated Teams bridge call or channel |
Hourly updates while Microsoft is executing |
| Stakeholder updates (business hours) |
Teams channel |
Every 2 hours |
| Out-of-hours incident |
On-call contact list (to be defined) |
As needed |
| Microsoft communication |
Severity A support ticket + direct Teams/email |
Minimum hourly during active restore |
5.2 Key Contacts (to be populated)
| Role |
Name |
Contact |
| OSDU Platform Team Lead |
TBD |
TBD |
| Task Lead |
TBD |
TBD |
| Microsoft Account Contact |
TBD |
TBD |
| Microsoft Support (Severity A) |
Azure Portal — support ticket |
aka.ms/azuresupport |
| Equinor Security On-Call |
TBD |
TBD |
| Data Producer Leads |
TBD |
TBD |
| Application Managers |
TBD |
TBD |
| Data Office |
TBD |
TBD |
| App Sec |
TBD |
TBD |
6. RTO / RPO Reference
| Metric |
Current State |
Notes |
| RPO (Recovery Point Objective) |
~24 hours |
Restore runs once daily at midnight UTC — data written after the last backup is lost |
| RTO (Recovery Time Objective) |
Target: 24 hours |
Actual in May 2026 test: ~48 hours (extended by Seismic subproject issues) |
| Time to detect |
Target: < 30 minutes |
Dependent on alert setup (see Section 1.2) |
| Time to declare |
Target: < 1 hour from detection |
Requires Task Lead availability — a deputy should be named |
| Time to begin restore |
Target: 2–3 hours from declaration |
Microsoft must be engaged quickly; USA time zone offset may delay execution |
| Time to complete restore |
~2–3 hours (proven in test) |
Microsoft-executed; exclude Seismic subproject complications |
| Time to validate |
Target: < 24 hours post-restore |
Extended in test due to Seismic bucket issue — mitigated going forward |
RPO acceptance required: The ~24-hour RPO must be formally accepted by the business. If a tighter RPO is required, this should be raised with Microsoft as a platform configuration change request.
7. Special Considerations
7.1 Seismic Store (Data Producers / DISKOS)
- The seismic subproject (
diskos) requires special handling — both the subproject metadata and the backing GCS bucket must be restored.
- Verify explicitly with Microsoft that the GCS bucket is included in the restore scope.
- ACL entitlement groups for the seismic subproject are not automatically restored — manual recreation via CLI may be required (process to be documented in runbook by OSDU Platform Team + Data Producers).
7.2 Security Response
- If the deletion event is suspected to be the result of a compromised service principal, credentials should be rotated before restore is initiated.
- The Equinor Security team should be looped in to assess the scope of any potential breach.
- Reference: OEP EntitlementsLogs and Azure AD sign-in logs are the primary sources for investigating suspicious SPN activity.
7.3 Data Partition Deletion
- Microsoft have confirmed that accidental data partition deletion is not currently recoverable (soft-delete is on the roadmap).
- If the
data partition itself is deleted, this is an L4 critical incident and senior leadership must be engaged immediately.
- This scenario requires a separate recovery runbook — to be developed with Microsoft.
7.4 Time Zone Considerations
- Microsoft restore execution is USA-based. Requests raised late in the European day may not be actioned until the following business day.
- For live incidents, use Severity A / Critical support ticket classification to ensure 24/7 response.
- Target to raise the Microsoft support ticket before midday CET where possible to avoid overnight delays.
7.5 Reingestion as an Alternative
- For small-scale deletions (L1–L2) where a full restore is disproportionate, reingestion of deleted records from source systems (Data Producers) is a viable alternative.
- This was validated in the May 2026 test (10 LogCurveType records reingested successfully).
- Data producer teams should confirm whether their source pipelines can be re-run selectively.
8. Open Items — Actions Required Before This Process Is Operational
| # |
Action |
Owner |
Priority |
| 1 |
Implement Azure Monitor alert for bulk DELETE operations on ADME |
OSDU Platform Team |
High |
| 2 |
Set up scheduled Log Analytics query for record count baseline monitoring |
OSDU Platform Team |
High |
| 3 |
Formally agree and document impact thresholds (Section 2) with the business |
OSDU Platform Team + Task Lead |
High |
| 4 |
Obtain formal RPO acceptance (~24 hours) from the business |
Task Lead |
High |
| 5 |
Name a Task Lead deputy for out-of-hours approval authority |
Task Lead |
High |
| 6 |
Populate key contacts table (Section 5.3) |
OSDU Platform Team Lead |
High |
| 7 |
Establish on-call rotation for OSDU Platform Team |
OSDU Platform Team Lead |
High |
| 8 |
Document Seismic Store ACL recreation process (with Data Producers) and add to runbook |
OSDU Platform Team + Data Producers |
High |
| 9 |
Agree communication cadence and support ticket process with Microsoft for live DR events |
OSDU Platform Team |
Medium |
| 10 |
Develop recovery runbook for accidental data partition deletion scenario |
OSDU Platform Team + Microsoft |
Medium |
| 11 |
Run a tabletop exercise against this process before the next DR test |
OSDU Platform Team |
Medium |
9. RACI Matrix
Role key:
| Code |
Role |
| PTL |
OSDU Platform Team Lead |
| PT |
OSDU Platform Team (engineers) |
| PO |
Task Lead |
| MS |
Microsoft (ADME Support) |
| SEC |
Equinor Security Team |
| DP |
Data Producers |
| APP |
Application Managers |
RACI key: R = Responsible · A = Accountable · C = Consulted · I = Informed
Phase 1 — Detection & Initial Assessment
| Activity |
PTL |
PT |
PO |
MS |
SEC |
Data Producers |
APP |
| Monitor alerts / identify anomalous deletion |
A |
R |
|
|
|
|
|
| Run record count checks against baseline |
A |
R |
|
|
|
|
|
| Review ADME audit logs & EntitlementsLogs |
A |
R |
|
|
|
|
|
| Classify incident level (L1–L4) |
R/A |
C |
I |
|
|
|
|
| Declare DR Incident |
R/A |
I |
I |
|
|
I |
I |
Phase 2 — Stakeholder Notification & Triage
| Activity |
PTL |
PT |
PO |
MS |
SEC |
Data Producers |
APP |
| Open Teams bridge / incident channel |
R/A |
C |
I |
|
|
I |
I |
| Raise Microsoft Severity A support ticket |
A |
R |
|
I |
|
|
|
| Notify Security team (if compromise suspected) |
R/A |
|
I |
|
I |
|
|
| Instruct data producers to pause ingest pipelines |
A |
R |
I |
|
|
I |
|
| Notify Application Managers of outage window |
R/A |
|
I |
|
|
|
I |
| Confirm backup state & restore point with Microsoft |
A |
R |
I |
C |
|
|
|
Phase 3 — Decision to Restore
| Activity |
PTL |
PT |
PO |
MS |
SEC |
Data Producers |
APP |
| Confirm scope and data loss from restore point delta |
A |
R |
C |
C |
|
|
|
| Obtain formal approval to proceed with restore |
C |
|
R/A |
|
|
|
|
| Document approval decision (scope, data loss, approver) |
A |
R |
I |
|
|
|
|
| Rotate SPN credentials (if compromise suspected) |
A |
R |
I |
|
C |
|
|
| Request double-confirmation from Microsoft before execution |
R/A |
C |
I |
I |
|
|
|
Phase 4 — Restore Execution
| Activity |
PTL |
PT |
PO |
MS |
SEC |
Data Producers |
APP |
| Initiate and execute restore |
C |
C |
I |
R/A |
|
|
|
| Verify seismic GCS bucket is included in restore scope |
A |
R |
|
C |
|
C |
|
| Monitor ADME service health during restore window |
A |
R |
|
C |
|
|
|
| Maintain hourly check-in with Microsoft during restore |
R/A |
C |
I |
R |
|
|
|
| Enforce write/ingest freeze on partition |
A |
R |
I |
|
|
I |
|
Phase 5 — Validation
| Activity |
PTL |
PT |
PO |
MS |
SEC |
Data Producers |
APP |
| Record count checks (all OSDU kinds vs. baseline) |
A |
R |
|
|
|
|
|
| ACL / entitlement group verification |
A |
R |
|
|
|
|
|
| Seismic subproject API & GCS bucket validation |
A |
R |
|
C |
|
R |
|
| Seismic ACL recreation (if required) |
A |
R |
|
|
|
C |
|
| Data Producers data validation |
C |
I |
|
|
|
R/A |
|
| End-to-end application manager validation |
I |
C |
A |
|
|
|
R |
| Formal incident closure sign-off |
R |
I |
A |
|
|
I |
I |
Phase 6 — Post-Incident
| Activity |
PTL |
PT |
PO |
MS |
SEC |
Data Producers |
APP |
| Produce incident report (timeline, root cause, data loss) |
A |
R |
I |
C |
C |
I |
I |
| Update DR runbook with gaps discovered |
A |
R |
|
C |
|
C |
|
| Update Seismic Store runbook section |
A |
R |
|
|
|
C |
|
| Conduct lessons learnt retrospective |
R/A |
C |
C |
C |
|
C |
C |
| Raise backlog items in ADO / Jira |
A |
R |
C |
|
|
|
|
| Review and update impact thresholds if required |
A |
C |
R |
|
|
C |
|
This document should be reviewed and updated following each DR test or live incident. Version history to be maintained alongside the DR runbook.
Last update:
2026-08-17