Claude Code is sold on what it can do. This post is about where it goes wrong when the work is not code. We built an entire course with it, seven sections of lectures, scripts and working files, and we kept a written record of the failures along the way. What follows comes from those post-mortems, not from a list of theoretical risks. Some of it is uncomfortable, because several of the mistakes were confident, plausible and wrong.
We are not the only ones asking the question. A developer piece this September set out what works and what does not for non-coding tasks. Our record adds the view from inside one long project, where the same failures came back until we changed how we worked.
"Instructions were not ignored deliberately. They were processed incompletely." Our own post-mortem, written after one section of the course
Where it was strong, so the rest is in proportion
It drafted at a volume we could not have matched by hand: lecture after lecture, each built from our source material, each in the agreed format. It restructured work across dozens of files when asked. Once a mistake was named precisely, it fixed it fast. The course exists because of that.
The failures below are not about whether it can write. They are about what happens between the instruction you gave and the result you get back.
The four ways it went wrong
| What happened | A real example from the course | The habit that catches it |
|---|---|---|
| Stated specifics it had never checked | A lecture named the wrong permission modes. Another said the tool runs in "three places", then later "four places". Neither count came from a source | Every specific fact, a name, a number, a setting, gets checked against a source document, or marked to check |
| Its own review agreed with its mistakes | An AI review in the same conversation passed lectures containing those errors, because it shared the knowledge the errors came from | Check facts against the source, not against a second opinion from the same tool in the same session |
| Long lists of instructions were partly done | One message held about twelve changes. Some were made, some missed, some made in one file but not the others | Number the instructions, and have each one confirmed by number before anything is called done |
| Filled a gap with a guess | Asked to set up two specific folders, it installed an entire skills library nobody had asked for | State the boundary explicitly, and ask it to stop and ask when an instruction could mean two different things |
"If it gets facts wrong, why use it at all?"
A tool that invents the names of its own settings is not a tool I want writing anything my team reads.
Because the errors were not random. They clustered in one place: specific facts it filled in from general knowledge rather than from a document in front of it. The structure, the drafting, the restructuring and the format were reliable. The details that needed a source were not.
That suggests a split rather than a verdict. Let it draft and organise. Make yourself the one who checks every name, number and setting against the source before anything goes out. It is the same split between assembly and judgment that our post on AI pilots describes, applied to facts.
"Can't I just ask it to double-check its own work?"
Surely the fix is to ask it to review what it wrote before I read it.
We did exactly that, and it was one of our six documented failures. The review ran in the same conversation as the writing. It had seen the same session and drew on the same knowledge, so it judged whether the lecture was internally consistent, not whether it was true. When the error came from that knowledge in the first place, the review agreed with it.
A self-review is useful for clarity, structure and tone. It is not a fact check.
"I'll give it everything in one message so it has the full picture."
It is better to explain the whole job at once than to drip-feed instructions.
Context helps. A long list of changes in one paragraph is a different thing, and it was the source of our most frustrating failure. Of roughly twelve instructions in one message, some were done, some were missed, and some were done in one file and left in others. Nothing was refused. The work simply came back incomplete, and it came back described as finished.
The fix is the numbered list from the box above. Same message, same context, but every instruction has a number, and the reply has to account for each one.
The part worth keeping
The tool did not fail by misunderstanding English or by refusing work. It failed between the instruction and the result: specifics it never checked, reviews that shared its blind spots, lists it half finished, and gaps it filled with a confident guess. Each of those has a habit that catches it, and none of the habits needs anything but the next message you send.


