Skip to main content
Applied AI for Management

What Claude Code Gets Wrong in Non-Coding Work: Lessons From Building a Whole Course With It

We built a seven-section course with Claude Code and kept a record of the failures. Facts it never checked, reviews that agreed with its mistakes, long instruction lists half done, gaps filled with guesses. Each one has a habit that catches it.

An AI assembles a tower of blocks, some cracked, beside a mirror reflecting the same stack and boxes marked with question marks, while a person reviews a checklist nearby.

Claude Code is sold on what it can do. This post is about where it goes wrong when the work is not code. We built an entire course with it, seven sections of lectures, scripts and working files, and we kept a written record of the failures along the way. What follows comes from those post-mortems, not from a list of theoretical risks. Some of it is uncomfortable, because several of the mistakes were confident, plausible and wrong.

We are not the only ones asking the question. A developer piece this September set out what works and what does not for non-coding tasks. Our record adds the view from inside one long project, where the same failures came back until we changed how we worked.

"Instructions were not ignored deliberately. They were processed incompletely." Our own post-mortem, written after one section of the course

Where it was strong, so the rest is in proportion

It drafted at a volume we could not have matched by hand: lecture after lecture, each built from our source material, each in the agreed format. It restructured work across dozens of files when asked. Once a mistake was named precisely, it fixed it fast. The course exists because of that.

The failures below are not about whether it can write. They are about what happens between the instruction you gave and the result you get back.

The four ways it went wrong

What happenedA real example from the courseThe habit that catches it
Stated specifics it had never checkedA lecture named the wrong permission modes. Another said the tool runs in "three places", then later "four places". Neither count came from a sourceEvery specific fact, a name, a number, a setting, gets checked against a source document, or marked to check
Its own review agreed with its mistakesAn AI review in the same conversation passed lectures containing those errors, because it shared the knowledge the errors came fromCheck facts against the source, not against a second opinion from the same tool in the same session
Long lists of instructions were partly doneOne message held about twelve changes. Some were made, some missed, some made in one file but not the othersNumber the instructions, and have each one confirmed by number before anything is called done
Filled a gap with a guessAsked to set up two specific folders, it installed an entire skills library nobody had asked forState the boundary explicitly, and ask it to stop and ask when an instruction could mean two different things

"If it gets facts wrong, why use it at all?"

A tool that invents the names of its own settings is not a tool I want writing anything my team reads.

Because the errors were not random. They clustered in one place: specific facts it filled in from general knowledge rather than from a document in front of it. The structure, the drafting, the restructuring and the format were reliable. The details that needed a source were not.

That suggests a split rather than a verdict. Let it draft and organise. Make yourself the one who checks every name, number and setting against the source before anything goes out. It is the same split between assembly and judgment that our post on AI pilots describes, applied to facts.

"Can't I just ask it to double-check its own work?"

Surely the fix is to ask it to review what it wrote before I read it.

We did exactly that, and it was one of our six documented failures. The review ran in the same conversation as the writing. It had seen the same session and drew on the same knowledge, so it judged whether the lecture was internally consistent, not whether it was true. When the error came from that knowledge in the first place, the review agreed with it.

A self-review is useful for clarity, structure and tone. It is not a fact check.

"I'll give it everything in one message so it has the full picture."

It is better to explain the whole job at once than to drip-feed instructions.

Context helps. A long list of changes in one paragraph is a different thing, and it was the source of our most frustrating failure. Of roughly twelve instructions in one message, some were done, some were missed, and some were done in one file and left in others. Nothing was refused. The work simply came back incomplete, and it came back described as finished.

The fix is the numbered list from the box above. Same message, same context, but every instruction has a number, and the reply has to account for each one.

The part worth keeping

The tool did not fail by misunderstanding English or by refusing work. It failed between the instruction and the result: specifics it never checked, reviews that shared its blind spots, lists it half finished, and gaps it filled with a confident guess. Each of those has a habit that catches it, and none of the habits needs anything but the next message you send.

AI Systems with Claude, for Scrum Masters and Project Managers

Most AI pilots in project management do not fail on the model. They fail because somebody pointed the thing at a decision instead of at a task. This course is built around that distinction, and around the work that follows once you get it right.

Seven sections, taught against one running project from the first lecture to the last. You watch a system get built, then you build the same one against your own sprint. Nothing here is a demo that works only on the example.

You finish with five working systems and you keep them. The sprint report machine turns your board export and standup notes into the report you currently write by hand. The retro intelligence system tells you what the team keeps saying, not just what it said this time. The stakeholder comms engine drafts the update in the register the audience expects. The backlog health monitor flags the quiet decay nobody has time to check for. The risk and dependency tracker follows the chains that actually bite.

Seventy-six working files ship with it, across four sections, so every system is built against real sprint data rather than material invented for a slide. Four sprints of retrospectives, board exports, standup notes, backlog and risk data.

Explore the Course


Build the Sprint Reporting System, Not Another Prompt

AI Systems with Claude is our own course for Scrum Masters and Agile Project Managers, taught against one running project across seven sections. You build five working systems and keep them: the sprint report machine, the retro intelligence system, the stakeholder comms engine, the backlog health monitor, and the risk and dependency tracker. Direct from HK School of Management, not on Udemy.

See What Is Included