
Outcome Before the Prompt
Write done-when before the first line of generation
The first prompt is the wrong place to invent the finish line.
If you open Codex — or Claude Code, Cursor, or any coding agent — and start typing before you know what “done” means, you are renting generation without a contract. The agent will produce something. You will review something. Whether that something was the job is left to tomorrow-you and a long scroll.
Write the outcome first. Then prompt.
The session outcome card
Before the first prompt, fill a short card. Keep it next to the work — top of the PR draft, a sticky note in REVIEW.md, or the first block in the thread.

- Goal — one sentence for what this session is for
- Done when — the acceptance check you can run or click
- Out of scope — what stays frozen this run
- Verify — one command, URL, or click path
- Rollback — how to undo in one step if this is wrong
If you cannot fill done-when, you are not ready to generate. You are still shaping the ask. That is fine — do it without burning a run.
Why the prompt is a bad place to decide
A first prompt mixes intent, constraints, and vibes. The agent optimizes for the loudest part of the paragraph. “Make auth nicer and also fix the flaky test” is two outcomes sharing one thread. Done-when collapses that into a check you can fail.
Done-when is not a feature list. It is a pass/fail. “Login returns 200 with a session cookie on a clean DB” is a done-when. “Improve the login experience” is a mood.
Out of scope is the other half. Name the polish you will not chase. If the agent offers it later, you already decided: new run, new card.
How to use the card in the session
Paste the card into the project instructions or the first message. Tell the agent: stop when done-when is met; do not expand out-of-scope; leave verify and rollback filled at the end.
When the run stops, you do not reread the chat for the thesis. You run verify. You check the files against out of scope. You call ship, kill, or one question against the card you wrote before generation started.
If verify fails, the failure is the run — not your memory of the prompt. If done-when was fuzzy, the failure was the card. Fix the card, then restart. Do not “nudge” a fuzzy outcome with ten more prompts.
What counts as done
Done means the acceptance check passed and the decision landed. A green verify step plus a one-line ship note counts. A long agent summary that never names the check does not.
Track for a week how often you start a run with a filled card versus a blank one. Compare rework: runs that came back after a ship because done-when was invented mid-thread. You do not need a dashboard — a tally in a note is enough.
Where it usually breaks
Prompt-as-spec: the first message is three paragraphs of product design. Compress it into the five fields. The prompt becomes “execute this card,” not “figure out what I meant.”
Moving done-when: every reply redefines success. Freeze the card when generation starts. Mid-run changes are a new session with a new card.
Verify you cannot run: “looks good in the chat” is not verify. If you cannot name a command or click path, you do not have done-when yet.
You skip the card when you are “just exploring.” Exploration is a valid goal — write it down: “Done when: three options ranked with tradeoffs, no code.” Then the explore run can end cleanly.
A few sharp edges
This is not anti-prompt craft. Prompt craft still matters. It just comes after the contract. A sharp prompt against a blank outcome is still a blank outcome.
Big work does not need a longer card. It needs a smaller done-when. Split the outcome before the first prompt. The card stays short; the tree of cards grows.
If you already keep durable asks with a receipt — goal and outcome without scrolling the harness — the card is that receipt. You do not need a broker to write one.
Generation is cheap. Ambiguous finish lines are not. Fill the card. Then open the prompt box.