Skip to main content
Applied AI for Management

Scrum Masters and Agile Project Managers: Why Most AI Pilots Die in the Same Place

Most AI pilots run by Scrum Masters and Agile PMs do not fail on the model. They fail because somebody pointed the thing at a decision instead of at a task. One test tells the two apart, and you can run it on your own backlog before you finish reading this sentence.

A project manager at a desk, seen from behind, facing a translucent dashboard. Handwritten sticky notes on the right and a kanban board on the left both flow into a central glowing icon, which outputs charts and status indicators

Here is the part that never makes it into the case study. Most AI pilots in project management do not fail on the model. They fail at one specific point, and it is the same point almost every time: somebody pointed the thing at a decision instead of at a task. Gartner expects more than 40% of agentic AI projects to be cancelled before the end of 2027, and found that only 28% of AI use cases in IT infrastructure and operations fully meet their ROI target, with one in five failing outright. Those numbers are not a verdict on the technology. They are a verdict on scoping. This post assumes you already run sprints and already have at least one reporting obligation you resent, because the fix is smaller than the failure rate suggests, and it is sitting in your own notes.

The clearest description of the failure we have found comes from Ardent Workshop, and it is worth reading twice. Autonomous project management agents, they write:

They work in tightly constrained demo environments and fall apart the moment they meet a real project with ambiguous status, political dynamics, and humans who don't update their tickets on time.

Read that again and notice what it does not say. It does not say the agent was not clever enough. It says the agent met a real project. Ambiguous status. Political dynamics. People who do not update tickets. None of those are technical problems, and none of them get better with a bigger model.

What happened when we built this

When we built the sprint reporting system that our course teaches, the first version read the Linear board. That seemed obvious. The board is structured, the notes are not, and structured data is what software wants.

It produced a sprint summary that was accurate and useless.

Every blocker it found was a ticket that had not moved. So it reported three blockers in a sprint that had exactly one, because two of the three were tickets somebody had finished on Thursday and not dragged across until Monday. And it missed the real blocker completely, because the real blocker was a person who would not commit to a date in writing, and there is no field on a board for that.

So we pointed it at the standup notes instead. Five text files, typed during the meeting, in the words people actually used. Worse grammar, no structure, no schema, nothing a database would recognise as data.

The summary it produced was the one we would have written by hand.

The notes worked because they contained the thing the board was missing. Not status. What people said. "Sam thinks the API change is going to bite us next sprint" is not a field anywhere, and it is the single most useful sentence in the week.

"Our data is too messy for that to work"

Maybe that works for a tidy team. Our board is a mess, half the tickets are stale, and nobody writes proper descriptions.

Good. That is the right starting condition, and the pilot that died last quarter probably died because somebody assumed the opposite.

The thing that decides whether AI works on a task is not how clean your data is. It is whether the task has a right answer. Here is the test.

The right-answer test. Three steps, and you can run it on the last thing you were asked to automate before you finish this paragraph.

  1. Write the task down as a sentence that starts with a verb. "Summarise what each person committed to in standup." Not "improve reporting." A verb and an object.
  2. Ask this question: if two competent people did this task from the same inputs, would they produce the same output? That is the whole test. Not similar output. The same output.
  3. If yes, it is assembly. Automate it, and expect it to work. If no, it is judgment. You are building a draft generator, not an automation, and you should scope it as one from the start.

Run it on a few real tasks and the line falls in a place that surprises most people.

"Collect what each person said they would do" is assembly. Two competent people produce the same list. "Write the sprint summary paragraph" is assembly, mostly, if the facts are fixed. "Decide whether the sprint goal is at risk" is judgment, and it is not close. Two competent Scrum Masters will disagree about that in the same sprint, with the same board in front of them, and both of them will be defensible.

Point an agent at the third one and you get a confident answer that is sometimes wrong, which is worse than no answer, because you cannot tell which kind you received.

Try it on whatever you were about to pilot. If it comes out as judgment, that is genuinely useful information, and nothing further down this post will change it.

"I am not technical enough to build this"

Fine, but the people who do this are developers. I run ceremonies. I am not going to learn to code to get my Friday afternoon back.

You do not need to, and the reason is that the hard part is not the code. The hard part is the input, and the input is five filenames.

The failure mode for a non-developer is not the technical step. It is starting from a blank page, with a blinking cursor and no structure, and trying to describe a whole system in one go. Start from a layout instead.

The five file week. This is the input side of the sprint reporting system, and it works with no tool at all.

  1. One plain text file per working day. monday.txt, tuesday.txt, wednesday.txt, thursday.txt, friday.txt. Nothing else in the names.
  2. All five in one folder, named for the sprint. We use Documents\sprint-notes\sprint-14. The sprint number in the folder name is doing real work, because it is what lets you compare this sprint to the last one later.
  3. Type during standup, in the words people use. Do not tidy it. Do not convert it into status. "Riley is blocked on the staging env, again" is better input than "Riley: blocked."
  4. Nothing else. No template, no headings, no fields.

That is it, and the naming convention is the mechanism rather than an administrative detail. Because the filenames carry the order, the whole week is readable in one pass, by a person or by anything else. There is no parsing step, no date extraction, nothing to configure. Five files in the right order is a week.

When we ran this on a real sprint, the Monday file was eleven lines long and contained the sentence that turned out to matter on Thursday. Nobody had logged it anywhere else.

If you already keep notes like this, you have the input side, and you can skip to the next section.

"My employer will not approve the tool"

Even if I could build it, we have a policy. I cannot put sprint data into some external tool and I am not going to ask.

That objection is usually correct, and it is also usually aimed at the wrong half of the system. Split it.

The two-layer split. Four steps, and the last one is the test that matters.

  1. Layer one is the inputs. Your notes, your filenames, your folder, on your machine. No tool, no account, no approval.
  2. Layer two is the assembly. Whatever reads those inputs and produces the summary. This is the layer your employer's policy actually governs.
  3. Build layer one first, and build it to stand alone. It has to be useful before layer two exists.
  4. Then run this test: if your employer banned every AI tool tomorrow, would layer one still be worth keeping? If yes, you built it in the right order. If no, you built a dependency and called it a system.

Layer one passes that test easily, which is the point. Five named text files per sprint is searchable, greppable, diffable against last sprint, and already better than a board for the one thing a board cannot hold. It survives any policy decision because it is not a tool. It is a filing convention.

The part worth keeping

The pilots that die are not the ambitious ones. They are the badly sorted ones. Somebody took a task that needed a judgment, handed it to something that produces confident output, and then discovered the problem in front of stakeholders.

The split is not sophisticated. Assembly or judgment, decided by asking whether two competent people would produce the same answer. But it is the thing that determines whether the next pilot survives contact with a real sprint, with its ambiguous status and its unmoved tickets and the one person who will not put a date in writing.

Your notes are better input than your board. Not because they are tidier, they are not, but because they contain what people said, and what people said is where the judgment calls are hiding.

The five file week is the input side, and you can run it on Monday without asking anyone for permission.

Sources: Gartner, via The Digital Project Manager, "AI Adoption in Project Management" | Ardent Workshop, "AI in Project Management 2026"

AI Systems with Claude, for Scrum Masters and Project Managers

Most AI pilots in project management do not fail on the model. They fail because somebody pointed the thing at a decision instead of at a task. This course is built around that distinction, and around the work that follows once you get it right.

Seven sections, taught against one running project from the first lecture to the last. You watch a system get built, then you build the same one against your own sprint. Nothing here is a demo that works only on the example.

You finish with five working systems and you keep them. The sprint report machine turns your board export and standup notes into the report you currently write by hand. The retro intelligence system tells you what the team keeps saying, not just what it said this time. The stakeholder comms engine drafts the update in the register the audience expects. The backlog health monitor flags the quiet decay nobody has time to check for. The risk and dependency tracker follows the chains that actually bite.

Seventy-six working files ship with it, across four sections, so every system is built against real sprint data rather than material invented for a slide. Four sprints of retrospectives, board exports, standup notes, backlog and risk data.

Explore the Course


Build the Sprint Reporting System, Not Another Prompt

AI Systems with Claude is our own course for Scrum Masters and Agile Project Managers, taught against one running project across seven sections. You build five working systems and keep them: the sprint report machine, the retro intelligence system, the stakeholder comms engine, the backlog health monitor, and the risk and dependency tracker. Direct from HK School of Management, not on Udemy.

See What Is Included