If you are learning AI right now, you are probably learning to prompt. Learn this first instead: before you hand any task to a model, write down what finished means. It is the skill almost everyone skips, and it is the one that will still matter after every prompt you know has expired.
At Edge8, my company, I watch founders and CTOs make the same mistake, and students learning AI are picking up the same habit. They write the goal carefully. They add some guardrails. Sometimes they set a budget. The model comes back with something that reads well, is nicely formatted, and answers the question. Then someone downstream finds the hole. And the hole is always in the one place nobody wrote anything down.
AI Stops at the Plausible Point, Not the Right One
Here is the mechanism you need to understand. When you don't tell a model what finished looks like, it doesn't refuse. It doesn't ask. It declares itself finished at the most plausible stopping point: the place where the output looks like outputs of that kind usually look.
Plausible is not the same as right. And plausible is exactly what a busy reviewer is scanning for, which is why these mistakes get through.
The cleanest example I have is one I wrote about on the blog at Edge8, my company. A PR agency in Australia runs a 90-day plan for every client, and it handed the plan to a model. It came back with plans that all started on July 1.
July 1 is the start of the fiscal year in Australia. If nobody tells you otherwise, it is the most plausible start date for a 90-day plan. But the agency's plans don't start on a fiscal quarter. They start on the contract date. For a client who signed on February 1, a plan starting July 1 means five months of nothing.
The goal was written. The guardrails were written. The budget was fine. Nobody wrote the line that said the plan begins on the signing date, so the model reached for the date plans most often begin on and called the job complete.
The model did not fail. The brief was never finished.
Why Students Skip It
Deciding when work is done feels like judgment. Judgment feels like the hard, valuable part, the thing you are asking the AI to supply. So it feels efficient, even respectful, to leave the finish line to the machine.
That is the mistake. A definition of done is not judgment. It is evidence. It is a short list of things that must be true about the output, written so that someone who did not do the work can read each line and answer yes or no. Judgment is how the model gets there. Evidence is how you know it arrived.
When you skip it, you are not handing the model a hard decision. You are handing it no decision at all.
A Number Is Not a Finish Line
The usual objection: "I already do this. I put numbers in my prompts. Ten slides. Two pages. Twelve quotes."
Another example from Edge8: a brief asked a model for 8 to 12 supporting quotes. It came back with 43. The range was right there in the brief. The model returned more than three times the ceiling anyway, because more felt more helpful, and nothing said what finished meant once it had twelve.
A number is a target. It tells the model what to aim at. A definition of done tells the reviewer what to check, including what happens to quote thirteen. It gets cut. Or the quotes get ranked and the top twelve survive. Or the extras get listed separately as overflow. Any of those works. Silence doesn't, because silence gets filled with whatever looks most plausible, and 43 quotes looks very thorough.
Weak vs Strong: One Task, Two Versions
Let's make this concrete with a task you could easily hand to AI this week: gathering supporting quotes for an argument, say for an essay, a project report or a pitch. This is an illustrative example, built on the research-brief case above.
The weak version
Weak: "Find 8 to 12 good, credible quotes that support my argument. Make sure they're relevant and well sourced."
Every word in that brief is an adjective or a target. "Good." "Credible." "Relevant." "Well sourced." None of them can be checked yes or no by someone else. This is the brief that produced 43 quotes.
The strong version
Same goal, plus a definition of done in five lines:
- The output contains no fewer than 8 and no more than 12 quotes. Anything beyond 12 is cut, not appended, not footnoted.
- Each quote is attributed to a named source with a link the reviewer can open.
- Each quote is mapped to the specific claim in my argument it supports. Two quotes for the same claim are reduced to the stronger one.
- No quote is paraphrased. If the exact wording can't be found, the source is listed as unverified instead of quoted.
- Reviewer counts the quotes first, then opens two links at random.
Now 43 quotes fails on line one in seconds. So does a beautifully organized set of twelve where three links go nowhere. You didn't need to be an expert to catch either. You needed the lines.
Run the July 1 plan through the same discipline and you get a first line like this: "Day one of the plan is February 1, the contract date, and the plan runs 90 days from that date, not from any calendar or fiscal quarter." Add one more line, "If the contract date is missing, stop and ask instead of picking one," and the July 1 output fails before anyone reads page two.
The Five-Line Checklist You Can Reuse
Five lines is the whole discipline. If your task needs more than five, the task is too big. Split it. Every line has to pass four rules:
- A stranger can check it. Someone who didn't write the brief and didn't do the work can read the line, read the output, and say yes or no.
- It names evidence, not adjectives. "Clear" and "thorough" are not evidence. "Every date is counted from February 1" is.
- At least one line says what must not be there. Models add. A list that only names what to include will be met and then exceeded.
- One line tells the reviewer where to look first. The first check should catch the most plausible wrong answer.
Test It in Five Minutes
Writing the lines is the easy part. Before you send the brief, run three quick tests:
- The stranger test. Give your five lines and a finished output to someone who never saw the brief. Can they answer yes or no on every line without asking you anything? If not, that line is an adjective in disguise. Rewrite it.
- The plausible-stop test. Read your brief as a smart, hurried intern would. Where would they stop and call it done? Run that stopping point through your lines. If it passes, your definition is too loose. The agency's brief passed the July 1 plan. Yours shouldn't.
- The wrong-answer test. Write down the single most plausible wrong output. A plan that starts on a fiscal quarter. Forty-three quotes. A report that ends with "it depends." Which line catches it? If none does, add one. If you can't think of a plausible wrong answer, you don't understand the task well enough to delegate it yet.
Those tests add about ten minutes to the brief. Skipping them costs the same ten minutes anyway, plus the afternoon of rework, plus the trust you lose when the hole shows up in front of someone who matters.
Why This Outlasts Prompting
A prompt is a workaround for what a particular model can't do yet. As I wrote in Your Prompts Are Expiring, every model generation quietly retires some of those workarounds. The clever instructions you memorize this year may be dead weight next year.
A definition of done is different. It describes your task, not the model. The plan still has to start on the contract date. The quotes still have to be real and linkable. The report still has to end in a recommendation. None of that changes when a new model ships. If anything, it gets more valuable, because stronger models finish faster, and finishing fast at the wrong point is still wrong.
It is also the skill that separates someone who uses AI from someone others trust to run work through it. Anyone can get a model to produce something that looks finished. Far fewer people can say, in advance and in writing, what finished means. That is the core of delegating to AI well, and it scales: the same five lines work whether you hand the task to one model, five agents or a human teammate.
Your Assignment This Week
Pick one task you are about to hand to AI. An essay outline, a research summary, a plan, anything. Before you write the prompt, write the definition of done in five lines. Run the stranger test on it. Then send it.
One task. Five lines. Before the prompt.
If you want to practice this with other people learning the same way, join the community at aiolabz.com. And if you want to prove you can do it, not just say you can, that is what AI Officer certification is for. Every challenge runs on your own data and workflows, and the certificate is proof of work, not proof of attendance.