Skip to content

Post-Incident Review (PIR)

ADME / OSDU Platform — Equinor

Document Owner OSDU Platform Team
Version 1.0
Created August 2026
Related Process Major Incident Response Process
Status Template

How to Use This Template

Complete this document within 5 business days of an incident being formally closed. It should be completed collaboratively by the incident lead and the OSDU Platform Team, with input from all stakeholders involved in the response.

The goal is not to assign blame — it is to understand what happened, why it happened, and what can be done to prevent it or respond better in the future.


Incident Summary

Field Value
Incident ID (e.g. INC-2026-001)
Title (brief description)
Date / Time Detected
Date / Time Declared
Date / Time Resolved
Total Duration
Environment(s) Affected (prod / test / dev)
Severity (P1 / P2 / P3 / P4)
Incident Type (Data Loss / Platform Availability / Security Breach / Infrastructure Failure / Dependency Failure)
Incident Lead
Participants

Incident Timeline

Provide a chronological account of events from first detection to resolution. Be specific about times and actions.

Time (UTC) Event Performed By Notes
First signal detected or report received
Initial investigation started
Incident declared
Stakeholders notified
Root cause identified
Remediation started
Remediation completed
Ticket severity downgraded (issue resolved; PIR ongoing — see note below) OSDU Platform Team Lead
Validation completed
PIR completed and actions raised
Incident closed

Note on severity downgrade: Once the platform issue is resolved and normal service is restored, the incident ticket severity should be downgraded to reflect that no active disruption remains. The ticket must stay open until the pre-closure checklist (Section 10) is fully complete. Downgrading must be agreed with the OSDU Platform Team Lead and communicated to all stakeholders.

Where did time get lost?

(Identify the phases where the response was slower than expected and briefly explain why.)


Root Cause Analysis

What Happened

(Describe what the platform state was before, during, and after the incident. What was the observable impact?)

Why It Happened — The Root Cause

(Describe the underlying technical or procedural cause. Use the "5 Whys" technique where helpful.)

Why Answer
Why did the incident occur?
Why did that happen?
Why did that happen?
Why did that happen?
Root cause:

Contributing Factors

(List any factors that made the incident worse or harder to detect/resolve — e.g. missing alerts, gaps in runbook, unclear ownership, tool limitations, timing.)

Factor Impact

What Prevented Faster Detection or Resolution?

(Was monitoring in place? Did alerts fire? Were the right people available? Was the runbook accurate?)


Impact Assessment

Area Detail
Data impacted (What data was affected? How much? Was any data permanently lost?)
Services impacted (Which ADME services were affected and for how long?)
Users / teams impacted (Which data producer teams or consuming applications were affected?)
Downstream impact (Any business processes or reporting affected?)
Duration of impact
Data loss window (if applicable) (e.g. data written between X and Y was lost)

What Went Well

(List things that worked effectively during the response — tools, communication, processes, individual actions.)

1. 2. 3.


What Didn't Go Well

(List gaps, delays, miscommunications, or failures in the response.)

# Issue Category Severity
1 (Procedural / Tooling / Communication / Access / Knowledge) (High / Medium / Low)
2
3

Lessons Learnt

(Summarise the key takeaways from this incident — insights that should inform how the team operates going forward.)

# Lesson Applies To
1 (Process / Tooling / Training / Documentation / Architecture)
2
3

Action Items — Prevention & Improvement

These are the concrete steps to be taken to prevent recurrence or improve the response capability. Each item must have a named owner and a target date.

# Action Category Owner Priority Target Date Status
1 (Prevention / Detection / Response / Documentation / Training) High / Medium / Low Open
2
3

Action categories:

Category Description
Prevention Architectural or process changes that eliminate or reduce the likelihood of recurrence
Detection Improvements to monitoring, alerting, or baseline tracking to catch similar events faster
Response Runbook updates, access changes, or process improvements to resolve faster next time
Documentation Missing or inaccurate documentation that contributed to delays
Training Knowledge gaps that should be addressed through training, walkthroughs, or rehearsals

Runbook & Process Updates Required

List any specific updates needed to the Major Incident Response Process, DR Trigger & Response Process, or other runbooks as a direct result of this incident.

Document Section Change Required Owner
Major Incident Response Process
DR Trigger & Response Process
Other:

Pre-Closure Checklist

All items below must be completed before the incident ticket can be formally closed. The OSDU Platform Team Lead is responsible for confirming each item.

# Task Completed Completed By Date
1 Platform service health confirmed — all affected services operating normally ☐
2 Record counts / data integrity verified against baseline ☐
3 Data producer teams have confirmed their data is accessible and intact ☐
4 Application Managers have confirmed end-to-end access ☐
5 All ingest pipelines (if paused) have been safely resumed ☐
6 Microsoft support ticket closed or downgraded (if applicable) ☐
7 Root cause identified and documented in this PIR ☐
8 Action items raised in ADO / Jira with owners and target dates ☐
9 Runbook / process update items identified and assigned ☐
10 Retrospective scheduled (P1/P2) or lessons noted (P3/P4) ☐
11 Incident report / PIR shared with Task Lead for sign-off ☐
12 Stakeholders notified of closure and provided with summary ☐

Follow-Up Review

Item Detail
Retrospective scheduled (Date / time — required for P1/P2; recommended for P3/P4)
Attendees
Action item review date (Date to check progress on items from Section 7)
PIR sign-off (Task Lead / Platform Team Lead)
Sign-off date

This document should be stored alongside the incident record and linked from the relevant ADO / Jira ticket. Action items should be tracked in the team backlog.


Last update: 2026-09-17