Go through every well-documented attack on an industrial plant over the past decade and ask one narrow question. What actually stopped it?
The answer is almost never a security product. It was operators driving to substations and closing breakers by hand. Clean offline backups, and staff who could run the plant on paper. An alarm, and a switch to manual. In one case, a bug in the attackers' own code.
That last one is worth sitting with. In 2017 a petrochemical plant in Saudi Arabia tripped itself to a safe shutdown twice. The investigation found malware built to reprogram the Triconex controllers that make up its safety system, the layer whose only job is to stop the process before people get hurt. FireEye assessed with moderate confidence that the objective was the capability to cause a physical consequence. It failed because the attackers' own deployment script had a defect. It worked, then backed itself out, which FireEye's analysts said they did not believe was supposed to happen.
The plant was saved by a mistake in the attack and by a safety system doing its day job. That is a strange thing to build a security programme around. It is also the most honest starting point available, and it should change what you fund first.
What actually contained the last decade of OT attacks
Take the incidents that have real public documentation, and a shape emerges.
In December 2015, attackers took down power for roughly 225,000 Ukrainian customers. They got in with stolen credentials over a VPN into the ICS network that had no second factor, then worked the breakers through remote HMI access much as an operator would. Power came back because operators drove to substations and closed breakers by hand. The attackers had anticipated exactly that and overwrote firmware on serial-to-Ethernet converters to block digital re-control. It still came back manually.
In March 2019, LockerGoga ransomware reached Norsk Hydro's production environment. This one matters because it is routinely miscited: the attack did spread from IT into production. Hydro moved plants to manual operation, staff worked on paper, and the company restored from clean isolated backups without paying. The bill was 60 to 71 million dollars, of which insurance covered about six percent.
In November 2023, an Iranian-linked group reached an internet-exposed Unitronics controller at a water booster station in Aliquippa, Pennsylvania, using default credentials. An alarm fired, staff took the system offline and ran the station manually. No boil-water advisory, no impact on drinking water.
Four incidents, four different sectors. Manual restoration, tested offline backups plus manual operation, an alarm and a manual fallback, and a bug in the malware. Not one of them was contained by a detection platform, a firewall rule, or a segmentation project.
The Oldsmar water case belongs here too, with a caveat. The 2021 account had someone remotely raise a sodium hydroxide setpoint from 100 to 11,100 parts per million while an operator watched the cursor move and reversed it. In 2023 a former city manager said a four-month FBI investigation could not confirm a targeted intrusion, and called it a non-event. No agency has published a retraction, so it sits unresolved. The lesson holds whichever version is true: a hardware dosing limit that caps the setpoint regardless of what the HMI says defends against a hostile operator and a tired one equally well.
The problem is not that operators do not know the list. It is the order.
Everyone in this field can recite the controls. Asset inventory, segmentation, secure remote access, monitoring, vulnerability management, incident response, backups, training. The eight-item version of that list is content marketing; the underlying items are not controversial.
Dale Peterson, who founded the S4 conference and sells no security products, has the sharpest explanation for why those lists keep getting longer. You are more likely to be blamed for leaving something off a cyber hygiene list, he argues, than for putting too much on it. Nobody ever got criticised for a ninth item.
He also counted what is on them. By his tally, roughly nine in ten items in both CISA's goals and IEC 62443's system requirements aim at reducing the likelihood of an attack rather than its consequences. The lists are overwhelmingly about keeping attackers out, in a field where the documented saves all came from surviving them getting in.
What the survey data shows is that organisations work the list in close to the wrong order. The SANS 2025 State of ICS/OT Security survey covers 330 practitioners. It is independently authored, though distributed by Dragos and Claroty, which sell several of the controls it measures. In it, asset visibility and inventory was the single biggest investment priority for 2026 and 2027, named by 54 percent. Threat detection came next at 43 percent, vulnerability management at 41.
Now put that against how attacks actually arrive. In the same survey, 22 percent of respondents had an ICS/OT incident in the previous twelve months, and half of those incidents came through external connectivity or remote access. Coverage of the controls addressing that exact pathway is thin: session recording at 13 percent fully implemented, ICS-aware device and protocol controls at 11 percent, real-time session approvals at 8 percent. Half the incidents, roughly one in eight organisations with the control. That is the widest gap in the dataset.
There is exactly one published ordering with an argument behind it
An unranked list of eight tells you nothing about what to do in a year when you can fund two. That is what makes most control lists useless under a budget constraint.
The exception is the SANS Five ICS Cybersecurity Critical Controls, published by Robert M. Lee and Tim Conway in late 2022. It is deliberately five, and deliberately ordered: ICS incident response, defensible architecture, ICS network visibility monitoring, secure remote access, and risk-based vulnerability management. The paper describes them as intelligence-driven, chosen from analysis of real compromises rather than derived from a framework.
Two steps in that order come with an explicit published argument. Take the second one first. Secure remote access sits at four, after visibility and monitoring at two and three, and the reason the paper gives is Colonial Pipeline.
Colonial was breached in May 2021 through a single legacy VPN account that was no longer meant to be in use. It had single-factor authentication, and its password had appeared in a prior breach dump. CEO Joseph Blount told the Senate it was a complicated password, not a Colonial123-type password. He was right, and it did not matter. The failure was not password strength. It was that nobody knew the account was still reachable.
The SANS argument follows directly: you cannot secure connectivity you cannot see, so the inventory of remote access paths has to precede the hardening of them. Thirty-one percent of organisations in the 2025 survey keep no formal inventory of their remote access points at all.
The reason incident response ranks first is stated too, and it is not the obvious one. Response planning establishes a shared view of the risks the organisation actually cares about, and that view then drives the requirements for everything else. Put response last and you get a set of controls chosen without reference to how you would ever use them. It also pays off on events with no attacker in them, because the same capability root-causes ordinary failures.
One more thing about Colonial, because it gets cited backwards. The ransomware never reached the OT network, but the pipeline stopped anyway, as a precautionary business decision made because the company could not confirm where the malware had and had not gone. That is evidence they could not tell. It is not evidence that segmentation worked.
The paper draws a general lesson from it that belongs in a project plan more than a security plan. When a new system goes in, both stay live through cut-over, because nobody can afford for the new one to fail. The old path is the one that never gets decommissioned. Decommissioning is the step that gets dropped when a migration runs late, and it is a security control.
Two honest caveats about the framework I just recommended
The second caveat is a plain gap. The five controls contain no backup requirement. Recovery appears only inside incident response. Adopt the five and nothing else, and you have no explicit instruction to back up controller logic, which is exactly what Hydro's recovery depended on. NIST SP 800-82 covers it, requiring backups of OT state, data, configuration files and programs, with restoration that puts human and environmental safety before restarting production. Take that part from NIST.
Your detection is thinnest exactly where the consequences are worst
The SANS 2025 survey asked about detection coverage by Purdue level, and the answer should reorder most monitoring roadmaps.
Full detection coverage runs at 9 percent at Level 4, 8 percent at Level 3, 7 percent at Level 2, and 3 to 4 percent at Levels 0 and 1. Levels 0 and 1 are the sensors, actuators and controllers, the layer where a compromise stops being a data problem and starts being a physical one.
SANS put it plainly in the report: coverage is sparse at best and concentrated far from where consequences are most severe, including remote field sites. Only 13 percent of organisations report full visibility across the ICS cyber kill chain.
The practical consequence is that a lot of OT monitoring is really IT monitoring that stops at the DMZ. It will catch the intrusion at Level 4, which is genuinely useful, and it will tell you nothing about ladder logic being rewritten down at Level 1.
Pushing IT tooling downward without adapting it has its own record, and the best-documented case is not a vendor anecdote. NASA's Office of Inspector General reported in 2017 that a security patch stopped monitoring equipment in a large engineering oven, starting a fire that destroyed the spacecraft hardware inside. The reboot the patch triggered also impeded alarm activation, so the fire burned undetected for three and a half hours.
Every number in this field comes from someone selling the fix
Here is the part that changes how you read the next list of eight.
CISA and NIST publish requirements for OT security. Neither publishes adoption statistics. Every widely-quoted number about what percentage of plants have monitoring, or segmentation, or an incident response plan, comes from a vendor or from a survey funded by vendors. That is not automatically disqualifying, but it does mean nobody in the measurement chain is neutral, and it shows.
Dragos publishes some of the most widely cited field data in the sector, drawn from its own customer engagements, and never discloses a sample size. Around its 2026 Year in Review, which covers 2025 data, the company published four incompatible versions of its own OT visibility figure. That 46 percent of assessments found adequate monitoring. That 46 percent found visibility gaps, which points the opposite way. That fewer than 10 percent of OT networks worldwide have meaningful monitoring. And that 30 percent have visibility. Those cannot all be true.
Dwell time is worse. SANS respondents say roughly half of incidents are detected within 24 hours. Dragos puts the industry average at 42 days. Both are defensible inside their own frames: survey respondents can only report incidents they detected, and Dragos only sees cases bad enough to trigger an external call. A reanalysis of the same SANS data by DeNexus, itself a risk-analytics vendor, computed a mean of 40.4 days, close to Dragos, and found detection getting worse year over year. Same underlying survey, opposite headline.
The counterweight nobody sells is also real. Waterfall's 2026 threat report counted 57 incidents worldwide with documented physical consequences in 2025. That is down about 25 percent, and the first decline in six years, even as ransomware volume against industrial organisations rose sharply. Waterfall counts only verified physical-consequence incidents, so treat it as a floor rather than a census. Both are true at once: more ransomware, fewer physical consequences.
And the one neutral baseline got quieter on OT. Comparing CISA's Cross-Sector Cybersecurity Performance Goals v1.0.1 from 2023 against v2.0 from December 2025, several OT-specific requirements soften. The monthly asset inventory cadence becomes "a more frequent basis." The deny-by-default rule for OT network connections becomes generic boundary language. The standalone OT training goal is folded into a general one. And the backup requirement, which in v1.0.1 explicitly listed programmable controller logic among the things that must be stored, no longer mentions it. That last one comes from comparing the two published documents rather than from any reporting, so check it yourself before relying on it. The phrase is in one version and absent from the other.
In fairness, the same revision cuts the other way, and it is the best evidence anyone has that prioritising works. CISA shortened its own list, from 37 goals to 34. The reason it gave was that the deleted items saw low adoption or overlapped with broader ones, and practitioners found them confusing. A government body measured its own checklist and pruned it, in the one place where nobody has a product to sell.
What to do first: cut the likelihood, then cut the consequence
If you have budget for two things this year, the evidence points somewhere unfashionable.
First, cut the likelihood. Build the inventory of every remote access path into your environment, including the vendor connections and the cellular modems nobody wrote down. Only then put multi-factor authentication and session brokering in front of the ones that survive review. That order matters, and it is the one step the SANS paper explicitly argues for: you cannot secure connectivity you cannot see. Note that this is an inventory of pathways, not of devices. It is a smaller and more finishable job than a full asset inventory, and it covers the route behind half the incidents organisations reported.
Second, cut the consequence. Rehearse running the plant without the network. Test restoring controller logic from backups you have actually opened. Confirm your operators can hold the process manually and know they are allowed to. In every one of the cases above, the thing that limited the damage was an operational capability, not a security tool.
The Saudi plant survived because a safety system did its job and an attacker made a mistake. You cannot buy the second one. You can absolutely build the first.
Sources: SANS 2025 State of ICS/OT Security (n=330) | SANS Five ICS Cybersecurity Critical Controls, Lee & Conway | NIST SP 800-82r3 | CISA CPG v1.0.1 and v2.0 | NASA Office of Inspector General, IG-17-011, February 2017 | Dragos 2026 OT Cybersecurity Year in Review (2025 data) | Waterfall 2026 OT Cyber Threat Report | FireEye/Mandiant TRITON analysis | Dale Peterson, S4x25 keynote, February 2025 | E-ISAC/SANS Ukraine analysis, 2016 | CISA AA23-335A | DeNexus | Senate testimony of Joseph Blount, June 2021

