ITSMAugust 24, 2026· 7 min read

Major Incident Management: What Happens in the Forty Minutes Before Anyone Declares One

There is a particular unfairness built into major incidents that everyone in IT operations knows but nobody highlights in the process documentation.

Most of them don't start with you. A payment provider goes down, a cloud region degrades, a vendor ships a bad update, and within the hour, it is your outage, your bridge call, and your name on the status updates. The cause was global, but the accountability is local and it is yours.

And these are the most visible hours of your year. A major incident is often the only time the executive floor watches IT work in real time. When a critical business function stops, the CEO gets interested, and the scrutiny that follows is predictable enough that PagerDuty has published a list of the thirteen questions executives ask after a major incident. Somebody, afterwards, will produce a list of things that could have been done better. The only real question is whether you wrote that list or received it.

So, the case for AI in major incident management is not really about speed, although the speed is real. It is about walking into that spotlight already knowing the scope, the timeline, who has been told what, and what happens next - instead of finding out on the call, in front of the people who fund your budget.

The place to look is the interval before anyone says the words "major incident." Everything that makes the response effective happens after that moment. Everything that makes it painful happens before it.

You set the detection rule yourself

First, how does the interval actually plays out. A payment gateway starts returning errors. The first ticket arrives and looks ordinary. So does the second, filed by someone in another region who describes it differently. The third comes from a monitoring alert, the fourth from a customer-facing team, the fifth from someone who assumes it is their laptop. Five people have now reported the same thing in five vocabularies, and none of them know about the other four. Somewhere in that interval a human connects the dots — too often because a senior stakeholder asked a question first.

The most important thing about automating that detection is that you set the rule, not the model.

In Rezolve.ai, major incident detection is configured rather than trained. An administrator defines what a cluster is: how many similar incidents, arriving inside what rolling window, judged similar in which fields — subject, assigned group, and others you choose. In the demo configuration, that is two similar incidents within two hours, but the point is that it is your number, written down, auditable, and changeable when your environment changes.

You also decide who declares. There is a toggle for automatic reporting, and it is off by default: the AI proposes, a human reviews and reports. Declaring a major incident has real cost — it pulls people out of other work and starts a clock the business can see — so that judgement stays with the service desk.

What the fifth person to report it sees

Here is the first place where the interval collapses.

When someone opens the fifth ticket about the payment gateway, the ticket itself tells them what is happening: this issue appears to belong to a cluster, four other tickets reported within three minutes, critical severity, five tickets in total. One click opens the cluster and shows the linked incidents, a written summary of what is happening across all of them, and the event trend.

The agent working on that ticket no longer needs to have seen the other four. The connection is made in the queue at the moment of contact rather than in someone's head afterwards.

One dashboard for emerging clusters

The second place it collapses is the dashboard.

Emerging clusters are surfaced centrally, with a state, a ticket count, and a growth rate. For e.g. five incidents in three minutes, growing, along with a recommendation. Open major incidents sit underneath linked incident counts, risk scores and age. Two buttons: report as a major incident, or report as a problem. Both are one click, and the distinction matters, because a cluster arriving in three minutes is an outage while the same cluster spread over three weeks is a problem record.

That single screen replaces the thing most major incident commanders actually do in the first twenty minutes, which is read tickets looking for a pattern they already suspect exists.

What happens when you declare

When the incident is declared, several things happen that nobody has to remember to do.

Every agent in the queue is notified. The related incidents are grouped and linked to the major incident record, so the five tickets become one piece of work with five pieces of evidence attached. The record itself is populated with a summary, the business impact and the affected services - written from what the tickets actually said, not from a template someone fills in later.

The war room opens with the context already in it

Most tools notify. That is the easy part, and it is not the same as coordinating.

Declaration creates a war room channel in Microsoft Teams, with the incident card, the summary and a link back to the record already in it. The responders who join do not arrive at an empty channel and a link because the context is already there, which removes the first ten minutes of every war room, the part where three people paste the same information to each other.

The assistant sits in the channel and can be tagged like a colleague. Ask it which incidents are linked, and it answers from the record rather than from someone's memory.

The loop almost nobody builds

This is the part worth reading twice, because it is where a major incident stops being purely a cost.

From the war room, the incident can be marked as a known issue. Once it is, the virtual agent tells every employee who starts a conversation that the outage exists. This is before they finish typing their question, with a start time and an ongoing status.

Think about what that does to the volume curve. During a major incident, most of what arrives at the service desk is the same incident reported again by people who cannot see each other. Those tickets consume agent attention at exactly the moment when attention is scarcest, and every one of them ends with a person waiting for an answer nobody has time to give them.

Closing that loop means the outage suppresses its own duplicate volume. The people affected get told rather than queued. The responders stay on the fault rather than on the inbox.

Detection shortens the interval before declaration. This shortens everything after it.

After the incident: the transcript lands on the record

When it is over, the war room transcript and a summary of it are captured back on the major incident record.

That matters more than it sounds. The post-incident review is the artefact everyone agrees is valuable and nobody wants to assemble at two in the morning, so it gets written days later from memory and partial screenshots, or it does not get written at all. A record that already contains the timeline, the linked evidence and the conversation is a review that starts from facts.

It is also the input to problem management. The same clustering that produced the major incident is what identifies the recurring fault underneath it — which is why the dashboard offers "report as a problem" next to "report as a major incident." One is what you do about the outage. The other is what you do about the cause.

What to ask any vendor, including us

Three questions separate real capability from a demo.

Show me the interval. From the first ticket to the moment a major incident is declared — what is that number today, and what does the product change about it? Not resolution time. The gap before anyone knows.

Show me what the twentieth person to report it sees. If the answer is a queue position, the product handles responders but not the people affected, and most of your ticket volume during an outage comes from the people affected.

Show me the configuration. If the clustering rules are not visible and editable — how many tickets, in what window, similar on which fields — then the threshold is someone else's judgement applied to your environment.

The thing that actually changed

Major incident management was designed around a constraint that no longer holds: a human had to notice the pattern before anything could begin. Everything else in practice like the accelerated procedure, the commander, the war room, the comms, was built to compress the time right after the incident notice. Nothing could compress the time before it, because there was nothing to build with.

That constraint has lifted, and it changes what prepared looks like.

The next major incident is coming, and it will probably start somewhere outside your walls again. What you control is the state you're in when the questions start: whether the scope was known before the third executive asked, whether the people affected were told before they filed tickets, and whether the review that follows is built from a record or from memory.

Those are the hours your year gets judged on. It is worth being equipped for them.

See the agentic service desk in action

Watch Rezolve.ai autonomously resolve real IT and HR tickets: governed, auditable, glass-box.

Book a demo

Frequently asked questions

What is a major incident?

An incident with serious business impact, handled under a separate accelerated procedure rather than the normal queue. The severity label is not what makes it major. The separate procedure is, because coordinating a response at scale is a different job from fixing a fault.

What is the major incident management process?

Declare, mobilize, communicate, resolve, close, review. The declaration step is the one most organizations handle informally, and it is where most of the elapsed time sits. Mature practices define the trigger quantitatively — how many tickets, of what kind, inside what window — so the declaration does not depend on someone noticing.

What does a major incident manager do?

They own the response rather than the fix. That means declaring the incident, assembling the right responders, running the bridge or war room, keeping stakeholders and affected users informed, deciding when service is restored, and owning the post-incident review. The technical resolution stays with the specialists.

How does ITIL 4 treat major incident management?

ITIL 4 covers it within the incident management practice and expects organizations to define major incidents separately, with their own criteria, roles, and procedure. It does not prescribe the thresholds, which is why the definition has to be written locally.

What is the difference between a major incident and a problem?

A major incident is an outage happening now, and the work is restoring service. A problem is the underlying cause, and the work is removing it so the disruption stops recurring. The same cluster of related tickets can produce either, depending on how fast they arrive: twenty in twenty minutes is a major incident, twenty over three weeks is a problem.

Can AI declare a major incident on its own?

It can, but in most environments it should not. Declaring a major incident pulls people off other work and starts a clock the business can see. In Rezolve.ai, automatic reporting is a toggle that is off by default: the system detects the cluster and proposes the declaration, and the service desk confirms it.

Paras Sachan

Get service-desk AI insights in your inbox

Practical guidance on agentic AI for IT and HR support: one email, no spam.

By submitting, you agree we may use the details you’ve provided to contact you. See our Privacy Policy.