HKSM › Books › ICS/OT Incident Response Drills and Exercises › The OT IR Lifecycle and the Plan Built for OT

Chapter 2

The OT IR Lifecycle and the Plan Built for OT

The Reflex That Works Everywhere Except Here

Every experienced IT incident responder carries the same reflex. When a machine is compromised, you take it off the network. Pull the cable, disable the port, quarantine the host, and investigate once the spread has stopped. For a laptop hit by ransomware, that reflex is exactly right, and a generic IT incident response plan built around it will carry you through the event. Now apply the same reflex to a compromised SCADA server, or to the engineering workstation that talks to the PLCs running a pressurized line. Disconnecting it may blind the operators at the worst possible moment, stop a physical process mid-cycle, or trip a safety interlock that turns a cyber problem into a hazardous one. The gap between an IT plan and an OT plan is therefore not a formatting difference or a matter of adding a few industrial contacts to the appendix. It is a difference in what can go wrong when the plan is followed. In IT, a bad response decision costs data, downtime, and money. In OT, it can damage equipment, injure people, or lose a process that takes days to restart. This chapter does two things. It takes the incident response lifecycle most teams already know and shows precisely where a physical process overrides the IT defaults. Then it sets out what an OT plan must contain that no IT template includes, so that when you start testing in later chapters, you are testing a plan that could actually work.

The Same Six Phases, Different Stakes

The incident response lifecycle that IT teams have used for decades applies to OT as well. It has six phases: Preparation, Identification, Containment, Eradication, Recovery, and Lessons Learned, and it is usually called PICERL: the first letter of each phase, in order. Each phase has a defined job. Preparation builds the capability before anything happens. Identification confirms that an incident is real and not a malfunction or a false alarm. Containment stops it from spreading. Eradication removes the threat. Recovery returns the environment to normal operation. Lessons Learned closes the loop by feeding what happened back into the next round of preparation. Keeping the same framework is useful, because it gives IT and OT teams a shared vocabulary, and in a real incident they will be working together under pressure. What changes is the execution. In IT, running these phases is largely procedural. In OT, three of the six phases need fundamentally different thinking, because the systems being contained, cleaned, and restored are connected to a physical process that keeps moving whether or not your response is ready. Those three are Containment, Eradication, and Recovery, highlighted in the figure below. Preparation and Identification also take on an OT shape, and Lessons Learned is the phase most organizations handle worst, so each phase gets attention here, but the three highlighted phases are where the IT reflex does the most damage.

The OT Incident Response Lifecycle

PICERL: six phases in a loop. Three of them need different thinking in OT.

  1. Preparation
  2. Identification
  3. ContainmentIsolate the network path first. Shut down only after safe state is confirmed. Pre-authorization required.
  4. Eradication
  5. RecoveryVerify controller logic against a known-good baseline before reconnecting.
  6. Lessons Learned

↺ Lessons Learned feeds back into Preparation: the lifecycle is a loop, not a line.

Standard IR phaseOT-distinct phase

Preparation and Identification Start on the Plant Floor

Preparation in OT includes three pieces of work that have no direct IT equivalent. The first is defining the safe state for each process area: the condition the process can be held in safely while response actions proceed. The second is training operators to recognize OT-specific indicators of compromise, the observable signals on the process and the control system that suggest unauthorized activity may be happening. The third is maintaining manual fallback procedures in a form operators can actually use under stress, not as a paragraph buried in an engineering document. Safe state and manual fallback are important enough to get their own chapter, Chapter 3, and the indicators operators look for are laid out layer by layer in Chapter 4. For now, the point is that all three must exist before an incident, because none of them can be produced during one. Identification looks different too. In IT, an incident often begins with an alert from a security tool. In OT, it usually begins with a process anomaly noticed by a person. An experienced operator sees a valve open without a command, a setpoint change that nobody authorized, or a historian trend that has simply stopped recording. That observation is the detection signal. It may turn out to be an instrument fault or a contractor's mistake, and part of Identification is working out which, but the plan has to treat the operator's observation as the beginning of the response, because very often nothing else will raise the alarm first.

Scenario: The Valve That Moved on Its Own

An operator on night shift watches a valve on the HMI change position. No command was issued from the console, no work order covers that unit, and the process trend shows the flow beginning to shift. An IT-trained responder joining the call remotely offers the familiar advice: if the HMI or the engineering workstation might be compromised, disconnect it now. Walk that advice forward. Disconnecting the HMI removes the operator's view of the process at the exact moment the process is behaving strangely. Shutting the unit down abruptly may trip interlocks and create a hazard of its own. Neither action is wrong because it is aggressive; it is wrong because it skips the questions the physical process forces you to answer first. The OT sequence runs differently. Isolate the network path the suspicious activity could be using, so that whatever is sending commands can no longer reach the controller. Assess whether the process can safely continue while you investigate. Move to shutdown only once the safe state has been confirmed. And every one of those steps should already be written in the plan, with the authority to take them agreed in advance, so the operator and the responder are following a decision made in daylight rather than negotiating one at 2 a.m.

Containment: Where IT Rules Fail

Containment is the phase where IT habits fail most visibly in OT. In IT, containment frequently means taking the affected system offline straight away, and the cost of doing so is measured in lost productivity. In OT, the same action may stop a physical process, trigger a safety interlock, or create a hazardous condition, and the cost is measured very differently. The correct OT sequence has three steps, and their order matters. First, isolate the network path: cut the route the attacker or the malware is using to reach the control system, while leaving the process and its local control running. Second, assess whether the process can safely continue in that isolated condition, a judgment that needs someone who understands the process and not only the threat. Third, shut down only after the safe state has been confirmed. This sequence must be pre-authorized in the incident response plan (IRP). If a containment action requires someone's sign-off, because it could stop production or affect safety, the plan should record who signs, under what conditions, and how they are reached, before any incident happens. A containment decision negotiated during an active incident is slower, more political, and more likely to be wrong than one agreed calmly in advance. When you start running exercises, containment authority is one of the first things worth testing, because it is where hesitation costs the most time.

Eradication and Recovery: Verify Before You Reconnect

Eradication in OT is not just removing malware from a set of Windows workstations. The systems that matter most, the controllers themselves, can be modified in ways no antivirus scan will find: changed logic, altered setpoints, an extra function block. So before any controller is reconnected, its logic has to be checked against a known-good baseline: a saved copy of the verified controller configuration, taken before the incident and stored where the attacker could not reach it. The comparison confirms that nothing was modified without authorization. Recovery follows from the same idea. In IT, recovery often means bringing a file server back from backup. In OT, it means reconstituting a control system from verified, isolated backups so that the process can be run under digital control again. The sequence is fixed: verify safety first, then verify configuration integrity, then reconnect under active monitoring so that any sign of reinfection is caught immediately. The temptation during a long outage is to skip the configuration check because the process is down and pressure from management is rising. That is the one shortcut that reliably makes things worse, because reconnecting a controller that still carries the attacker's changes reintroduces the very threat you were trying to eradicate. There is no shortcut on this sequence in OT. Whether your backups and baselines are actually good enough to support it is a separate question, and Chapter 10 is devoted to answering it honestly.

Lessons Learned, and Why the Lifecycle Is a Loop

Lessons Learned is the phase most organizations do last and do worst. A meeting is held after the incident or the exercise, observations are written down, the document is filed, and six months later the same gaps appear in the next exercise as if for the first time. The reason is almost always the same: the findings had no owners and no deadlines. A process that produces observations without assigning them produces nothing of lasting value. The output of this phase has to be a concrete improvement plan that names what specifically changes, who is responsible for making the change, and by what date. Every finding left without an owner is a gap you have agreed to rediscover. How to write that improvement plan in a form that survives contact with a busy plant is the subject of Chapter 9. What matters here is the shape of the lifecycle as a whole. PICERL is not six steps executed once and filed away. It is a loop, and each completed cycle feeds the Preparation phase of the next. Identification gaps sharpen the list of indicators operators are trained to watch. Containment failures become pre-authorized decisions in the next revision of the plan. Eradication problems update the baseline and verification procedures. Every drill and exercise in this book is a way of running that loop on purpose, without waiting for an attacker to start it for you, and each pass should leave the next incident less damaging than the last.

Safety Before Containment

With the lifecycle in place, turn to the document that has to carry it: the plan itself. The first principle of an OT incident response plan reverses the IT order of operations. The IT sequence is to contain the threat and then investigate. In OT, following that sequence blindly can get people killed, because containment actions act on systems that control physical equipment. Physical safety supersedes cyber containment, every time. The correct OT sequence is to ensure safety, establish a safe state, and then investigate. Each process area's safe state is a predefined, documented condition that the team can reach without touching the compromised system, so that reaching it does not depend on equipment that may no longer be trustworthy. Any containment action that could create a physical hazard requires explicit pre-authorization, and that authorization has to come from someone who understands both the cyber threat and the process consequence. In many organizations those two kinds of understanding sit in different departments, which is precisely why the decision cannot be left until the incident. A security lead may know that a compromised workstation must be isolated without knowing that isolating it removes the operators' only view of a reactor. A process engineer may know the reactor without being able to judge how far the compromise has spread. If your plan does not start from the safety-first sequence and record who makes these calls, it is not merely incomplete. It is dangerous, because it invites people to follow IT instincts in a place where those instincts can hurt someone.

Five Elements Your IT Template Is Missing

An OT incident response plan must define five elements that no generic IT template includes. Each exists because of a specific way OT incidents differ from IT incidents, and each tends to be discovered missing the first time a plan is tested. The first two, safe state and manual fallback, decide what happens to the physical process while the response is underway; they are defined briefly in the table below and treated fully in Chapter 3, because they need input from process engineering and operations before the security team can write anything. The remaining three, operator authorization limits, a pre-populated vendor contact list, and evidence preservation procedures, decide how the response itself runs in the first hours, and they are covered in the sections that follow. Read the table as an audit tool. For your own plan, ask whether each element exists, whether it is specific to your site and process rather than copied from a template, and whether the people who would use it have ever seen it. An element that exists only in the document has not yet passed any of those three tests.

Element What the plan must define Why an IT template misses it
Safe state A predefined condition for each process area that the team can reach without using the compromised system IT systems have no physical process that must be held safely while they are investigated
Manual fallback How operators run the process from local controls, bypassing digital systems An IT service that goes offline simply stops; a process has to keep being controlled by someone
Operator authorization limits Exactly what the first person on the scene may and may not do before the response team arrives In IT, the first responder is usually a security professional, not a plant operator
Vendor contact list Pre-populated emergency and after-hours contacts for every critical control system vendor and integrator OT recovery often depends on vendors the IT team has never dealt with
Evidence preservation Approved ways to collect evidence without stopping the process, and how chain of custody is kept IT forensics assumes you can image a disk; you usually cannot image a running controller

What the Operator Is Authorized to Do

The plant operator is usually the first person to notice an OT incident, and very often the first person to act on it. That makes the operator's authority one of the most consequential lines in the whole plan, and one of the most commonly left blank. The plan has to define exactly what that person may do before the security team arrives, and the permitted actions are deliberately few. Document the observation: what was seen, on which screen or instrument, and at what time. Avoid touching the affected systems, because well-meant actions such as rebooting an HMI or reloading a program can destroy evidence and, worse, can trigger the very process consequences an attacker was aiming for. Alert the shift supervisor, so that the decision moves to someone with authority to start the wider response. There is one exception, and the plan should state it plainly: on an immediate safety event, the operator executes the defined safe-state procedure without waiting for anyone. The reasoning behind such tight limits is not distrust of operators. It is that operators are not security professionals, and a plan cannot assume they will improvise well in a situation they have never trained for. If the plan does assume it, then the outcome of your incident depends on whoever happens to be on shift that night. How the supervisor's alert travels upward and when it becomes a formally declared incident is the subject of Chapter 4.

Vendor Contacts and Evidence You Can Actually Collect

Two more elements decide how much time the response loses in its first hours. The first is the vendor contact list. Searching for an emergency support number during an incident costs hours, and in OT those hours matter, because recovery of a DCS, a PLC platform, or a historian often depends on the vendor's engineers and the specific integrator who configured your system years ago. The plan should pre-populate emergency contacts for every critical vendor, covering at a minimum the DCS, the PLCs, the historian, any remote access provider, and the OEM integrators, each with after-hours numbers that someone has recently confirmed still work. The second element is evidence preservation, which in OT is more limited than IT responders expect. Pulling a hard drive or imaging the memory of a running controller is often impossible without stopping the process. The approved options are less dramatic but still valuable: network captures taken from a span port, photographs of HMI screens showing what operators saw, exports from the historian covering the period of the event, and preserved operator logs. Each piece of evidence also needs a chain of custody: the documented record of who collected it, when, and how it was handled afterwards, so that it remains usable for investigation, insurance, or legal action. Define that chain before the incident. People collecting evidence for the first time, at night and under pressure, will not invent a sound custody process on the spot.

Write It for the Operator at 2 A.M.

All of this leads to a simple test for any OT incident response plan, and it is worth applying to yours before you test anything else. The test is not whether the plan looks complete. It is whether an operator at 2 a.m., with a physical process running and the phones ringing, could pick the document up and follow it. That is a very different document from the one a security consultant drafts to satisfy a compliance audit. The audit version tends to be long, organized around frameworks, and written in the language of policy. The operator's version has to be short where it needs to be acted on, organized around the process areas the operator actually sees, and explicit about what to do first. In practice, that means building the plan in a particular order. Start from the safety-first sequence and the safe state for each process area. Define what the operator is authorized to do and, just as importantly, what they must not do. Pre-populate every vendor and internal contact so that nobody has to search. Document which evidence can be collected without stopping the process, and how to keep it. Then hand the draft to an operator and watch where they hesitate. Each hesitation is a finding, gathered at no cost before any drill has been scheduled.

What's Next

Two of the five elements carry more weight than the rest, because without them the plan's containment steps have nowhere to go. Chapter 3 takes safe state and manual fallback in depth: what they mean for continuous, batch, and discrete processes, who has to define them, and why the procedures behind "switch to manual" have to exist before anyone writes the security plan.

Reflect

  • If your current plan was written from an IT template, which step in its containment section would cause the most harm if an operator followed it literally on your control network?
  • Who at your site can authorize a containment action that might stop production, and is that authority written in the plan or only understood informally?
  • Do you have a known-good baseline for your critical controllers, and when was it last compared against what is actually running?
  • What happened to the findings from your last incident review or exercise? Can you name an owner and a date for each one?
  • If you handed your plan to a night-shift operator today, where would they hesitate first?

How To Land the Job and Interview for Project Managers Course

Take the next big step in your project management career with HK School of Management. Whether you're breaking into the field or aiming for your dream job, this course gives you the tools to stand out, impress in interviews, and secure the role you deserve.

This isn’t just another job-hunting guide—it’s a tailored roadmap for project managers. You’ll craft winning resumes, tackle tough interview questions, and plan your first 90 days with confidence. Our hands-on approach includes real-world examples, AI-powered resume hacks, and interactive exercises to sharpen your skills.

You'll navigate the hiring process like a pro, with expert insights on personal branding, salary negotiation, and career growth strategies. Plus, downloadable templates and step-by-step guidance ensure you're always prepared.

Learn from seasoned professionals and join a community of ambitious project managers. Ready to land your ideal job and thrive in your career? Enroll now and take control of your future!

Explore the Course


Stop Managing Admin. Start Leading the Future!

HK School of Management helps you learn AI prompt engineering for project work. Move beyond status reports and risk logs with practical prompt frameworks for everyday tasks. Practical skills, tools, and guidance you can apply right away. Covered by Udemy's 30-day refund policy.

Enroll Now