What to review when the AI checks its own work

What's left to review when the AI checks its own work

8 min read

Anthropic describes its newest model as "much stronger at verifying its work and iterating carefully until it succeeds." That is good news, and it is worth being precise about what kind.

It means the draft that lands on your desk has already been checked once. Numbers reconciled against the source. Structure compared to whatever you said done looks like. The obvious contradictions caught and fixed before you ever saw them. The floor rises, and a lot of the tedious first pass — the pass where you were mostly hunting for arithmetic slips — stops earning its keep.

It also means the draft arrives wearing a QA stamp it issued to itself.

We've written before that polish is camouflage: the most dangerous draft is the one that reads clean enough to skip checking. A model that verifies its own work produces drafts that read cleaner. The risk doesn't go away — it gets better dressed.

So the question isn't whether to keep reviewing. It's which parts of your review just became redundant, and which parts were never the model's to do.

The checks that genuinely transfer

Three kinds, and you should hand them over without guilt.

Traceability. Does every number in this draft appear somewhere in the inputs? Does every quote match the source verbatim? This is mechanical comparison — exactly what a machine should do, and it's the check most likely to catch the error that would embarrass you. Ask for it explicitly and you'll get it.

Internal consistency. Does the pricing in section four match the scope in section two? Does the timeline in the summary match the timeline in the body? Long documents drift; catching drift is pattern-matching across a text the model just wrote and still has entirely in view.

Compliance with the stated standard. If your brief said "under 300 words, five sections, every open item carries a date," the model can check that better than you can, because it doesn't get bored on the fourth read. This is the strongest argument for writing the done-test down: a criterion you wrote is a criterion the model can now audit itself against.

Between them, those three are most of what a first pass used to be. Reclaiming that time is the real gain here — bigger than any capability headline.

The checks that never transfer, and why

The remaining four aren't hard for a model in the sense that it lacks the ability. They're impossible in a stricter sense: the information required to make the judgment isn't in the workspace. No amount of self-verification reaches a fact the model was never given.

Did the client actually say this? A draft can cite a commitment that traces cleanly to your meeting notes — and your meeting notes can be wrong, or compressed, or a paraphrase you wrote at speed. The model verifies against the inputs. You verify the inputs against the room you were in. Only one of you was there.

Does this match the relationship? A recommendation can be correct and still wrong to send. The client whose engineering lead has a competing theory, the sponsor who needs to look good in front of their CEO, the account where "we recommend pausing" would read as an exit — none of that is in the folder, and most of it shouldn't be. This is the check that protects the engagement rather than the document.

This is the same discipline that decides what stays out of a weekly client update: the scope request you deflected verbally and don't want to concede in writing, the internal politics you'd only discuss on a call. A model handed those notes has no way to know that one line is a landmine and the next is routine — they look identical as text. Both are true, both trace to your inputs, and only one of them should reach the client. Verification cannot help here, because nothing about the sentence is wrong. It's the consequence that's wrong, and consequences live in the relationship rather than in the document.

Would you defend this number in a renegotiation? Traceable and defensible are different standards. A figure can trace perfectly to an input and still be one you'd rather not stake a renewal on. That call is about your exposure, and the model has none.

Does it sound like you? The one everyone underestimates. Voice failures don't look like errors — they look like competent writing by somebody else, which is precisely what a self-verifying model will pass as correct. Reading a paragraph aloud remains the fastest test anyone has invented.

Notice what these have in common. They aren't harder versions of the mechanical checks. They're a different category: judgment that depends on facts the workspace doesn't contain. That's why they'll still be yours after the next release, and the one after that.

What this looks like in a real review pass

Take a case-study brief from our own kit. Its review pass has five steps, and the split falls straight through the middle:

  1. Rewrite the first three sentences in your own wordsyours. Voice, and it's the part prospects actually read.
  2. Trace every number and every quote to an input filetransfers. Ask for it and it's done before you look.
  3. Do the non-numeric outcomes feel real, or padded?yours. "Feels padded" is a judgment about your standards.
  4. Did the anonymization level hold throughout?mostly transfers. Mechanical scan for names and identifying details, with your eye on whether a detail is identifying in context — an industry plus a headcount plus a city is often a name.
  5. Confirm permissions match what the client actually agreed tonever transfers. That agreement happened on a call. It isn't in the folder, and it shouldn't be.

That's roughly half the pass reclaimed and half sharpened. The half you keep is the half where being wrong is expensive.

The brief pattern that follows

Here's the move that turns a self-verifying model from a risk into an instrument. Add one line to the end of any brief you hand over:

Before you finish: check this draft against "done looks like" above, and list every criterion you could not verify from the inputs.

That second clause is the whole trick. You're not just asking for a self-check — you're asking for a map of the self-check's blind spots. What comes back is a short list of exactly the places where the model was working from something other than evidence, and that list is where your attention belongs.

It's a direct descendant of a rule we already build into every workspace: never invent a fact, write [NEEDED: what's missing] instead. The [NEEDED] convention marks gaps during drafting. This marks gaps after verification — the things the model filled in plausibly enough that it didn't flag them, but can't actually stand behind.

Two things make this work better than a general "be careful" instruction. It's checkable, so you can tell whether the model did it. And it scales inversely with risk: a clean, well-briefed run returns a short list, and a run built on a thin brief returns a long one — which tells you the brief was thin before the draft costs you anything.

What comes back

Concretely. A consultant runs the case-study brief on a wrapped engagement — a funnel diagnostic for a mid-size SaaS client — and appends the line. The draft arrives, and under it:

Checked against "done looks like": 600–900 words ✓ (740). Every factual claim traces to an input ✓. At least one non-numeric outcome alongside the numbers ✓. Anonymization level "named with industry context" held throughout ✓.

Could not verify from the inputs:

  • The kickoff notes describe the trial-to-paid rate as "about 9%." The final report states 9.1%. I used 9.1% as the more precise figure, but I can't confirm the two refer to the same cohort and date range.
  • The closing email says the team "finally stopped arguing about the funnel." I've written this as a stated outcome about internal alignment. Whether that's a durable change or a good week isn't in the inputs.
  • No input shows what the client agreed to regarding publication. I've assumed the anonymization level you specified is sufficient.

Read that list and notice what it did. It didn't find errors — the draft passed its own checks. It found the three places where the draft is load-bearing on something other than evidence, which is exactly where a case study gets you in trouble: a number that might be comparing two different cohorts, an outcome that might be a mood, and a permission nobody confirmed.

Every one of those is a five-second answer for you and an impossible one for the model. You know whether the 9% and the 9.1% are the same cohort, because you built the sheet. You know whether the arguing stopped for good, because you were on the calls. And you know exactly what was agreed about publishing, because you agreed it.

That's the review pass now: not hunting for mistakes in a document that already checked itself, but answering three questions only you can answer.

If your standing setup already carries the standards, this costs nothing to adopt. Our Claude Workspace Kit ships six engagement briefs that each define what done looks like and pair it with a review pass — the objective, the inputs checklist, the constraints, and the human check written out per deliverable type. Adding the could-not-verify line to briefs that already state their done-test is a one-line change with an outsized return. And if your "done" definitions live in documented procedures rather than in your head, that's what the Solo Operator's SOP Bundle is for: a standard is what turns "review the draft" from vibes into a checklist a model can be held to.

Want to feel the difference before changing anything? Our free Friday Update Brief ships with a sample week of messy notes and the output the brief should produce. Run it, then run it again with the could-not-verify line appended, and compare what comes back. On a complete brief the list should be nearly empty — which is itself the lesson.

The part worth keeping

Delegation was never abdication, and a model that checks its own work doesn't change that — it moves the boundary. The mechanical half of your review is genuinely going away, and you should let it. What's left is smaller, harder, and more clearly the thing you're actually paid for: knowing what was said in the room, what the client can hear, which numbers you'd defend, and whether it sounds like you.

That was always the valuable half. Now it's the only half.

Free, and complete

Run the Friday Update Brief on a real week.

A week of raw notes in, a client-ready update out. It ships with a sample week and the reference output, so there is something to compare against.

Get it free →