Chapter 3
Safe State and Manual Fallback
An aluminum smelter runs its potlines continuously, with the molten bath in each cell held at over 950 degrees Celsius. The cells cannot simply be switched off and restarted later: if they are allowed to cool, the aluminum solidifies inside them and the equipment is destroyed. One night a cyber attack takes the SCADA system offline. The screens the operators use to watch and adjust the cells go dark, and the shift supervisor, facing a room full of people waiting for direction, says the most natural thing in the world: "We will figure it out." It sounds like leadership. It is actually an admission that nobody ever decided what the potline should do when its digital control disappears, which condition it should be held in, who confirms that condition has been reached, and what the operators do with their hands in the meantime. Every minute spent figuring it out is a minute in which the process drifts without a plan, and the potline does not wait for the security team to finish a meeting. That is not an incident response plan. It is a $50 million equipment loss waiting to happen. Chapter 13 tells what happened when a real aluminum producer faced this situation, and how much of the outcome was decided before the attack began.
What Safe State Actually Means
Chapter 2 set out the first rule of an OT plan: ensure safety, establish a safe state, and only then investigate. That rule is only as good as the definition of safe state behind it, and in many plans the phrase appears without one. Safe state: a defined physical condition of the process that is safe to maintain while a cyber investigation is underway. Each part of that definition matters. It is defined, meaning written down in advance with specific values, not left as a general intention to "make things safe." It is physical, meaning it describes the state of the process and its equipment, such as a load level, a temperature range, or a set of valve positions, rather than the state of the network. And it is something that can be maintained for a period of time, because an investigation may take hours or days, and a condition that can only be held for twenty minutes does not buy the response team anything. The easiest way to find your safe state is to answer one question for each process area: if digital visibility is lost, what is the last known-good operating condition that operators can hold manually? That question rules out answers that depend on the compromised system still working, and it forces the conversation to the floor, where the people who will actually hold the condition can say whether it is achievable. It also explains why safe state is not something the security team decides on its own, which is the subject of the next section.
Who Defines It, and Why Not During the Incident
Safe state is not vague, not improvised, and not chosen by the security team. Responsibility splits three ways. Process engineers define it, because only they understand which conditions the equipment and the chemistry or physics of the process can tolerate. Security teams document it, placing it in the incident response plan alongside the cyber triggers that would invoke it. Operators execute it, because they are the people on shift when it is needed. And both engineering and security must agree on the definition before any incident, so that the plan and the process do not contradict each other. That three-way split is exactly why safe state cannot be worked out under pressure. Process engineers and operators are rarely in the same room when an incident is active: the engineers may be off site, asleep, or in a different time zone, while the operators are dealing with a process behaving in ways they do not understand. The conversation that defines safe state has to happen in advance, in daylight, with the right people present. The resulting procedure has to be written down where operators can find it. And operators must have actually run through it before they need it, because reading a procedure for the first time during an emergency is how steps get missed. An incident response team that arrives at a plant and asks "what is the safe state?" will usually be met with silence, or with several confident and contradictory answers, which is worse.
Safe State Looks Different by Process Type
There is no universal safe state, because the right condition depends on what the process does and how it responds to being interrupted. A useful starting point is to sort your process areas into three broad types. Continuous processes, such as power generation and refining, run around the clock and suffer when stopped abruptly, so their safe state is typically a hold at reduced load or a controlled partial shutdown that keeps the equipment within its limits. Batch processes, such as food and pharmaceutical production, move through defined runs, so their safe state is usually to complete the batch already in progress and halt any new starts until the investigation is finished. Discrete manufacturing, such as automotive assembly, can usually stop completely without physical consequence, because a stopped assembly line is a production loss rather than an equipment hazard. The stakes rise accordingly. In a discrete plant, choosing the wrong safe state costs output. In a continuous process, getting it wrong means equipment damage, which is why the potline at the start of this chapter is such an unforgiving example. Most real sites are not one type. A single facility may run a continuous core process, batch preparation upstream, and discrete packaging downstream, and each area needs its own definition. Use the table below to place each area before you write anything, and expect the continuous areas to need the most engineering time.
| Process type | Examples | Typical safe state | Cost of getting it wrong |
|---|---|---|---|
| Continuous | Power generation, refining | Hold at reduced load, or a controlled partial shutdown | Equipment damage |
| Batch | Food, pharmaceutical production | Complete the current batch, halt new starts | Lost or compromised batch |
| Discrete | Automotive assembly | Stop completely | Usually no physical consequence; lost production |
Writing a Safe State Definition
A safe state definition does not need to be long, but it does need two things that most draft definitions leave out. It must name the target condition in terms an operator can check, for example 40 percent load, rather than "reduced load" or "stable operation." And it must name the triggering criteria, the specific situation that tells the team to move to that condition, for example loss of DCS visibility. Without a target, operators will each hold the process wherever they personally judge to be safe. Without a trigger, nobody knows when the definition applies, and the decision drifts back into improvisation. The table below shows the fields a complete definition carries, using the two examples above. The ownership fields are as important as the technical ones. They record the agreement between process engineering and security described earlier, and they make it obvious when a definition has been written by one group without the other. The last field, the date operators last practiced the procedure, is the one most often blank. It is also the one that tells you whether the definition exists anywhere other than on paper, and it is a natural target for the short, single-procedure drill described in Chapter 1. Write one definition for each process area, starting with the area where the wrong answer would be most expensive.
| Field | What it records | Example |
|---|---|---|
| Process area | The unit or area this definition covers, and its process type | Main generating unit (continuous) |
| Target condition | The physical condition to reach and hold, stated in values operators can check | 40 percent load |
| Triggering criteria | The situation that tells the team to move to the safe state | Loss of DCS visibility |
| Defined by | The process engineer who set the condition | Process engineering, named individual |
| Documented by | The security owner who placed it in the incident response plan | OT security lead, named individual |
| Executed by | The role that carries it out on shift | Shift operators, under the shift supervisor |
| Last practiced | When operators last ran through the procedure | Often blank: the first gap to close |
Manual Fallback: What "Switch to Manual" Really Requires
Safe state tells you where the process should be held. Manual fallback tells you how operators hold it there when the digital systems cannot be trusted. Manual fallback: the documented set of procedures for operating a process without its digital control systems, using local controls in the field. Every OT incident response plan must include these procedures, and the reason becomes obvious as soon as you read a typical containment section. Many plans include "switch to manual operation" as a containment step, a sensible way to cut an attacker off from the process while keeping it running. But that sentence is only a step if a manual procedure stands behind it. If no such procedure exists, the plan effectively ends at that line. When the security lead says "go manual," the operators need a document that tells them exactly what that means for their unit: which local panels to man, which valves and breakers they will be operating by hand, which readings to take from field instruments now that the HMI is not trustworthy, how often to take them, and how many people that takes per shift. A plant that has run under automatic control for years may have very few people left who have ever operated it any other way, and the knowledge often sits with one or two senior operators. Writing it down is the only way to make sure it is still available on the night it is needed, whoever is on shift.
Fallback Comes Before the Security Plan
The order in which these documents are written matters more than it first appears. If the manual fallback procedures for your process do not exist, process engineering and operations must create them before anyone writes the security plan, not alongside it and certainly not after it. The reason is that the security plan depends on them. Its containment steps assume that the process can be run without the systems being isolated, its safe state definitions assume operators can hold a condition by hand, and its recovery steps assume the process was kept in a known condition while controllers were verified and restored. A security plan written first will contain sentences that point at procedures nobody has written, and as Chapter 1 showed, gaps of that kind stay invisible on review and surface only when the plan is run. Writing the fallback first also changes the security conversation for the better. Once engineering and operations have worked out how long a unit can realistically run in manual, how many people that needs, and which conditions cannot be held without automation, the security team knows how much time containment and recovery can actually take, and the plan can be built around real limits instead of hopeful ones. When you audit your own plan, follow every reference to manual operation back to its procedure. Each one that leads nowhere is a finding, and it belongs to process engineering and operations as much as to security.
Define It Before You Need It
Go back to the potline at the start of this chapter. At 950 degrees, it does not wait for the security team to finish a meeting, and the moment the SCADA system goes offline, somebody on that shift is already making a decision about what the process should do next. There is no version of the night in which nobody decides. The only question is whether that decision follows a written procedure, agreed in advance by the people who understand the process and practiced by the people who have to carry it out, or an improvised one made by whoever is in the room. Everything in this chapter is a way of moving that decision out of the incident and into the planning. Sort your process areas by type. For each one, write a safe state definition with a target condition and a triggering criterion, and record who defined it, who documented it, who executes it, and when it was last practiced. Follow every "switch to manual" in your plan back to a procedure that operators have actually seen, and if the procedure does not exist, ask process engineering and operations to write it before anything else is revised. None of this needs an attacker, a vendor, or a budget to get started. It needs the right people in a room for an afternoon. You will not have that afternoon during the incident, so take it now.
What's Next
A safe state is only useful if someone recognizes, early enough, that it is time to use it. Chapter 4 turns to detection: how operators notice that something on the process is wrong, the OT-specific indicators that point to a cyber cause, and the point at which an anomaly on one screen becomes a declared incident with the whole plan behind it.
Reflect
- For your most critical process area, what is the last known-good operating condition operators could hold by hand if digital visibility were lost tomorrow?
- Which of your process areas are continuous, which are batch, and which are discrete, and has anyone defined a different safe state for each?
- Find every place your plan says "switch to manual." Does each one lead to a written procedure that the current operators have actually read?
- Who at your site still knows how to run the process without automation, and what would happen to that knowledge if they retired next year?
- When did process engineering, operations, and security last sit in the same room to agree on what the process should do during a cyber incident?
AI for Agile Project Managers and Scrum Masters
Become an AI-first leader and transform your agile practice by leveraging artificial intelligence as your most powerful co-pilot. This course is designed to help you drive efficiency, insight, and innovation, ensuring you stay at the forefront of a rapidly evolving project management landscape.
This isn't about replacing human intuition—it's about augmenting it. You'll master prompt engineering to automate mundane tasks, freeing up your time for high-impact strategic leadership and creative problem-solving. Learn to refine backlogs, create strategic roadmaps, and integrate AI seamlessly into your agile ceremonies.
Gain predictive power by using AI-driven insights to anticipate project risks and seize new opportunities for more reliable outcomes. We deliver practical, prompt-based workflows and proven strategies built around real-world agile challenges that you can implement immediately within your framework.
Master foundational AI concepts specifically relevant to Scrum environments while developing advanced skills to handle diverse agile scenarios. You will learn to champion an AI-enabled culture within your organization, fostering a dynamic environment of continuous improvement and superior team delivery.
Learn alongside professionals around the world working in project leadership and delivery.
Explore the CourseLaunch your Agile career!
HK School of Management helps you learn Agile and Scrum with practical playbooks, AI-powered prompts, and real-world workflows for planning and delivery. Practical skills, tools, and guidance you can apply right away. Covered by Udemy's 30-day refund policy.
Learn More