Most AI demos show the output and stop. This post shows the output, the inputs it came from, and the four things it got wrong. The example is the first thing Claude Code builds in our course: a summary of one week of standup notes from a five-person team, produced from a one-sentence instruction. It reads well. It is also wrong in four places, and the way it is wrong is more useful than any polished demo, because the same two mistakes will show up in your own summaries.
This was the whole instruction, typed into Claude Code in a folder holding five standup notes, one per day:
"Read the standup notes in this folder and write a summary of what the team completed this week. Save it as sprint-summary.md." The instruction used in our course's first-build lecture
The inputs, and what came back
The notes cover one week for a fictional team of five building a customer dashboard: Alex, Sam, Riley, Chris and Taylor. Each day's file records what each person did yesterday, what they plan today, and any blockers. They are written the way real standup notes are written: plainly, in the words people used, and in the past tense about the previous day.
Claude Code read all five and wrote a summary. Its "In progress" section, unedited, said:
CSV export feature (Alex) - 80% done, edge cases being handled
Error state handling across dashboard (Chris) - 70% done
Performance testing full suite (Riley) - running, results being written up
Notification system end-to-end testing (Taylor) - testing in progress
And under "Blockers resolved this sprint", it said the activity feed and the sprint metrics widget were both "unblocked Wednesday when Riley published the API spec".
Here is what the notes actually say:
| The summary said | The notes say | Kind of error |
|---|---|---|
| CSV export is 80% done, in progress | Friday: "CSV export done and tested. Merged." | Stale status |
| Error state handling is 70% done, in progress | Friday: "Error state handling complete and tested." | Stale status |
| Activity feed unblocked Wednesday | Tuesday: Alex, "Yesterday - Got Riley's API spec." | Misdated event |
| Metrics widget unblocked Wednesday | Tuesday: Chris, "Got the API spec from Riley." | Misdated event |
Everything else checks out: the seven completed items, the two items genuinely still in progress, the blocker resolved by design approval, the two slow queries fixed on Thursday. The summary is mostly right, which is exactly why the four errors would survive a quick read.
Two mistakes, not four
The errors fall into two kinds, and both come from what the instruction did not say.
Stale status. Both items appear as "in progress" with the percentages from Thursday's note. Friday's note says they were finished. Nothing in the instruction said that the latest note is the final word, so the summary took the more detailed, earlier description.
Misdated events. Standup notes say "yesterday". Tuesday's note saying "yesterday I got the spec" means Monday. The summary dated the unblock to Wednesday, the day after the notes recorded it. Nothing in the instruction explained how the notes encode time.
"Four errors in one short summary. It isn't ready."
If it gets four things wrong in a week of notes, I am not putting my name on anything it writes.
You should not put your name on this summary as it stands, and that is the point of publishing it. But look at what kind of wrong it is. It did not invent work or misattribute anything to the wrong person. It made two systematic mistakes about time, both of which a person can catch in minutes with the notes open, and both of which can be written into the instruction.
Systematic errors are the good kind. They repeat, which means a rule can address them. Random errors are the ones that should worry you.
"If I have to check it, I might as well write it myself."
The whole point was to save me the time. Checking it line by line is the same job.
It is not quite the same job. Checking means reading a short summary with the notes beside it and asking three specific questions. Writing means building the summary from nothing. In this example, all four errors were found by comparing seven completed items and four in-progress items against five short files, one pass, one direction.
And the check gets shorter as the instruction improves. Every error you turn into a rule is one you stop having to look for.
"Our standups aren't written down anywhere."
We do standup on a call. Nobody takes notes. There is nothing to summarise.
Then there is no input, and no tool can summarise a conversation nobody recorded. The notes in this example are the system. They are ordinary, unpolished, written in the words people used, and that is what made the summary possible.
Starting them costs very little. A plain text file per day, typed during the standup you already attend, is enough. Our post on why AI pilots fail describes the five-file layout these notes follow.
The part worth keeping
A one-sentence instruction produced a summary that was mostly right and wrong in two predictable ways. Neither of them was the tool failing to understand English. Both were the instruction failing to say how standup notes work: that the last note is final, and that "yesterday" is relative.
That is the general lesson. When an AI summary is wrong in a pattern, the fix is usually a sentence in the instruction, not a better tool.


