Post-Incident Review (PIR)
ADME / OSDU Platform — Equinor
| Document Owner | OSDU Platform Team |
| Version | 1.0 |
| Created | August 2026 |
| Related Process | Major Incident Response Process |
| Status | Template |
How to Use This Template
Complete this document within 5 business days of an incident being formally closed. It should be completed collaboratively by the incident lead and the OSDU Platform Team, with input from all stakeholders involved in the response.
The goal is not to assign blame — it is to understand what happened, why it happened, and what can be done to prevent it or respond better in the future.
Incident Summary
| Field | Value |
|---|---|
| Incident ID | (e.g. INC-2026-001) |
| Title | (brief description) |
| Date / Time Detected | |
| Date / Time Declared | |
| Date / Time Resolved | |
| Total Duration | |
| Environment(s) Affected | (prod / test / dev) |
| Severity | (P1 / P2 / P3 / P4) |
| Incident Type | (Data Loss / Platform Availability / Security Breach / Infrastructure Failure / Dependency Failure) |
| Incident Lead | |
| Participants |
Incident Timeline
Provide a chronological account of events from first detection to resolution. Be specific about times and actions.
| Time (UTC) | Event | Performed By | Notes |
|---|---|---|---|
| First signal detected or report received | |||
| Initial investigation started | |||
| Incident declared | |||
| Stakeholders notified | |||
| Root cause identified | |||
| Remediation started | |||
| Remediation completed | |||
| Ticket severity downgraded (issue resolved; PIR ongoing — see note below) | OSDU Platform Team Lead | ||
| Validation completed | |||
| PIR completed and actions raised | |||
| Incident closed |
Note on severity downgrade: Once the platform issue is resolved and normal service is restored, the incident ticket severity should be downgraded to reflect that no active disruption remains. The ticket must stay open until the pre-closure checklist (Section 10) is fully complete. Downgrading must be agreed with the OSDU Platform Team Lead and communicated to all stakeholders.
Where did time get lost?
(Identify the phases where the response was slower than expected and briefly explain why.)
Root Cause Analysis
What Happened
(Describe what the platform state was before, during, and after the incident. What was the observable impact?)
Why It Happened — The Root Cause
(Describe the underlying technical or procedural cause. Use the "5 Whys" technique where helpful.)
| Why | Answer |
|---|---|
| Why did the incident occur? | |
| Why did that happen? | |
| Why did that happen? | |
| Why did that happen? | |
| Root cause: |
Contributing Factors
(List any factors that made the incident worse or harder to detect/resolve — e.g. missing alerts, gaps in runbook, unclear ownership, tool limitations, timing.)
| Factor | Impact |
|---|---|
What Prevented Faster Detection or Resolution?
(Was monitoring in place? Did alerts fire? Were the right people available? Was the runbook accurate?)
Impact Assessment
| Area | Detail |
|---|---|
| Data impacted | (What data was affected? How much? Was any data permanently lost?) |
| Services impacted | (Which ADME services were affected and for how long?) |
| Users / teams impacted | (Which data producer teams or consuming applications were affected?) |
| Downstream impact | (Any business processes or reporting affected?) |
| Duration of impact | |
| Data loss window (if applicable) | (e.g. data written between X and Y was lost) |
What Went Well
(List things that worked effectively during the response — tools, communication, processes, individual actions.)
1. 2. 3.
What Didn't Go Well
(List gaps, delays, miscommunications, or failures in the response.)
| # | Issue | Category | Severity |
|---|---|---|---|
| 1 | (Procedural / Tooling / Communication / Access / Knowledge) | (High / Medium / Low) | |
| 2 | |||
| 3 |
Lessons Learnt
(Summarise the key takeaways from this incident — insights that should inform how the team operates going forward.)
| # | Lesson | Applies To |
|---|---|---|
| 1 | (Process / Tooling / Training / Documentation / Architecture) | |
| 2 | ||
| 3 |
Action Items — Prevention & Improvement
These are the concrete steps to be taken to prevent recurrence or improve the response capability. Each item must have a named owner and a target date.
| # | Action | Category | Owner | Priority | Target Date | Status |
|---|---|---|---|---|---|---|
| 1 | (Prevention / Detection / Response / Documentation / Training) | High / Medium / Low | Open | |||
| 2 | ||||||
| 3 |
Action categories:
| Category | Description |
|---|---|
| Prevention | Architectural or process changes that eliminate or reduce the likelihood of recurrence |
| Detection | Improvements to monitoring, alerting, or baseline tracking to catch similar events faster |
| Response | Runbook updates, access changes, or process improvements to resolve faster next time |
| Documentation | Missing or inaccurate documentation that contributed to delays |
| Training | Knowledge gaps that should be addressed through training, walkthroughs, or rehearsals |
Runbook & Process Updates Required
List any specific updates needed to the Major Incident Response Process, DR Trigger & Response Process, or other runbooks as a direct result of this incident.
| Document | Section | Change Required | Owner |
|---|---|---|---|
| Major Incident Response Process | |||
| DR Trigger & Response Process | |||
| Other: |
Pre-Closure Checklist
All items below must be completed before the incident ticket can be formally closed. The OSDU Platform Team Lead is responsible for confirming each item.
| # | Task | Completed | Completed By | Date |
|---|---|---|---|---|
| 1 | Platform service health confirmed — all affected services operating normally | ☐ | ||
| 2 | Record counts / data integrity verified against baseline | ☐ | ||
| 3 | Data producer teams have confirmed their data is accessible and intact | ☐ | ||
| 4 | Application Managers have confirmed end-to-end access | ☐ | ||
| 5 | All ingest pipelines (if paused) have been safely resumed | ☐ | ||
| 6 | Microsoft support ticket closed or downgraded (if applicable) | ☐ | ||
| 7 | Root cause identified and documented in this PIR | ☐ | ||
| 8 | Action items raised in ADO / Jira with owners and target dates | ☐ | ||
| 9 | Runbook / process update items identified and assigned | ☐ | ||
| 10 | Retrospective scheduled (P1/P2) or lessons noted (P3/P4) | ☐ | ||
| 11 | Incident report / PIR shared with Task Lead for sign-off | ☐ | ||
| 12 | Stakeholders notified of closure and provided with summary | ☐ |
Follow-Up Review
| Item | Detail |
|---|---|
| Retrospective scheduled | (Date / time — required for P1/P2; recommended for P3/P4) |
| Attendees | |
| Action item review date | (Date to check progress on items from Section 7) |
| PIR sign-off | (Task Lead / Platform Team Lead) |
| Sign-off date |
This document should be stored alongside the incident record and linked from the relevant ADO / Jira ticket. Action items should be tracked in the team backlog.