The effort dial: how much thinking does a task deserve?

The effort dial: how much thinking does a task deserve?

6 min read

There's a control in Claude now that most consultants will ignore for months, and it's the one most likely to change how their week actually goes.

Claude Opus 5 has an adjustable effort setting. You dial how much work the model does before it answers — how much it thinks, checks, and reconsiders. Low effort answers fast. High effort takes longer and arrives more carefully. Anthropic frames it as trading cost against capability, which is accurate and slightly undersells what it means in practice.

Here's what it means in practice: the model-selection decision you've been making at the top of every task now has a second axis, and the second axis is more useful than the first for most consulting work.

Why this is the same decision you already know how to make

We've written before that the question which actually picks a model is how long the leash is: how much work you hand over between check-ins, and what wrong costs you. A subject-line rewrite is a ten-second leash — you see the output immediately, a miss costs nothing. A "here's the transcript, my services file, and my rate floor, draft the proposal" run is a long leash, where a failure in the middle compounds quietly until the end.

Effort is that same question with a knob instead of a menu. Short leash, cheap wrongness, you're reading every word anyway: low effort, and enjoy getting your afternoon back. Long leash, expensive wrongness, a draft you'd hate to be subtly wrong: turn it up and let the model do the checking before you ever see the output.

The reason this matters more than the model choice is arithmetic. Most consultants have maybe two models genuinely available to them on a given plan. Effort gives you four or five settings on every task, including the routine ones — which means it's a lever you'll pull twenty times a week rather than twice.

The tasks that earn maximum effort

Three shapes, and they're the ones you'd expect if you've been reading along.

Assembly across many sources. A case study built from a project folder — kickoff brief, three monthly updates, final deliverable, the closing email where the client described the change in their own words. The work is holding a dozen inputs coherent while writing to a standard. Effort buys you exactly the thing that fails first here: sustained attention through the middle.

Absorbing something long before answering. A forty-page RFP. A contract you need to actually understand rather than skim. The failure mode at low effort isn't a wrong answer, it's a confident answer built on the first eight pages.

Anything where you can't cheaply check the output. This is the real criterion, and it's worth stating plainly: if the cost of you not noticing an error is high, buy the extra thinking. A proposal with a plausible-but-off price, a case study citing a metric nobody produced — these are expensive precisely because they survive a quick read. Anthropic says Opus 5 is "much stronger at verifying its work and iterating carefully until it succeeds," and high effort is where that behaviour shows up. Let it check itself, then check it anyway — the parts of the review that survive a self-verifying model are a shorter list than you'd think, and a more expensive one.

The tasks that are actively worse at high effort

This is the half people skip, and it's where the wins are.

The Friday client update. Known shape, your inputs, your proofread before it sends. High effort buys you a longer wait for a draft you were always going to read line by line. The update's quality comes from its shape and its inputs, not from how hard the model thought.

Anything conversational. Talking a document into shape, arguing with yourself about positioning, "what am I missing here." You want the loop, not the leash. A better answer ninety seconds later is a worse thinking partner than a good answer now — the rhythm is the value.

Work where your judgment gates every stage anyway. Scoping something unsettled, pricing something strange. You're going to intervene after each step regardless, so the model's extra deliberation is deliberation you're about to override.

There's a subtler cost too. Effort spends your plan's usage allowance, and it spends it faster at the top of the range. The failure mode is the same one that used to happen with models: everything creeps up the dial out of habit, and by the end of the month you're rationing on the work that actually needed the headroom.

Put it in the same place as your model defaults

The practical move takes ten minutes. You may already have a list of your five most frequent Claude tasks with a model next to each — if you don't, that list is the highest-leverage ten minutes in this whole topic.

Add a second column.

Task Tier Effort
Email passes, quick reformats fast low
Friday client update from notes fast low
Discovery debrief from a messy transcript strong default medium–high
Scope and pricing you steer step by step strong default low–medium
Proposal draft from a full brief top you can reach high
Case study from a project folder top you can reach high
Long-document absorption (RFP, contract) top you can reach high
Thinking partner strong default low

The specific values matter less than having them. The win isn't the assignments — it's never re-deciding at a blank conversation again, which is the same win the model list gave you.

Two patterns worth noticing in that table. Everything routine sits at low effort on the fast tier, which is cheaper and faster than what most people are doing today. And the high-effort rows are all long-leash rows — the dial and the leash rule agree, which is what makes this easy to remember.

The thing effort cannot fix

A dial that controls how hard the model thinks does nothing about what the model knows.

If your handoff doesn't include the rate floor, maximum effort produces a beautifully reasoned proposal priced below your floor. If it doesn't say the client got burned by a big-bang project last year, you get a confident recommendation for a big-bang project. Effort improves the reasoning applied to the material you supplied; it cannot improve the material.

This is worth internalising because the dial is genuinely seductive. When output disappoints, turning it up feels like the obvious response, and it's usually the wrong one. The first question is always whether the brief was complete — objective, inputs, constraints, what done looks like, and why. In our experience the honest answer is "no" far more often than "the model didn't think hard enough."

Test that on yourself cheaply: our free Friday Update Brief is a complete handoff with a sample week of messy notes and the reference output it should produce. Run it at low effort and again at high. If the two are close — and on that task they should be — you've learned something more useful about your workflow than any benchmark could tell you: that the brief was doing the work, not the dial.

And if you'd rather your standing setup carry the constraints so every run starts complete, that's what the Claude Workspace Kit assembles: the instructions file with the guardrails written, four deliverable standards, and six ready-to-run briefs where the objective, inputs, and done-test are already specified. The dial decides how hard Claude thinks. The folder decides what it's thinking about.

Set your defaults once. Spend the dial where wrongness is expensive. And when the output disappoints, check the brief before you reach for the knob.

Free, and complete

Run the Friday Update Brief on a real week.

A week of raw notes in, a client-ready update out. It ships with a sample week and the reference output, so there is something to compare against.

Get it free →