For most of the past two years, the default move with a capable AI model has been to dial everything up — more reasoning, more depth, more tokens — on the assumption that harder problems deserve harder effort. That instinct is understandable but, it turns out, expensive in exactly the wrong places.
A more useful pattern has quietly emerged: start cheap, then spend precisely.
The two-step logic
The idea is straightforward. Use the model’s lowest effort setting to produce a draft artifact — a spreadsheet, a deck, a document. The draft does not need to be good enough to ship. It needs to be good enough to reveal where the real problems are.
That distinction matters because knowledge work, unlike code, has no built-in feedback mechanism. A spreadsheet does not tell you when its argument is wrong. A slide deck does not flag when the narrative falls apart between slides three and seven. The only way to find those gaps is to have something to look at. A low-effort draft produced in seconds gives you exactly that — a structure to interrogate, not a finished product to polish.
Once the draft exists, the expensive move becomes targeted. Ask one specific, hard question: which three assumptions most change the outcome? What evidence does each slide claim to show, and does it actually show it? What is this document asking the reader to do, and where does it bury the answer? That focused question runs at higher effort — or passed to a second model acting as auditor, not collaborator — because it requires genuine judgment about what matters.
The final revision then drops back to low effort, with an explicit instruction not to rebuild the whole file. Without that instruction, models have a tendency to silently replace everything they touch — a helpful instinct that destroys the work you just reviewed.
Why this holds up for spreadsheets, decks, and documents
The pattern applies differently by artifact type, and those differences are worth understanding.
For spreadsheets, the useful question after the draft is not “does this look right” but “what are we actually assuming.” A working model mixes reported facts, transaction terms, outside estimates, and the builder’s own assumptions without labeling them. Separating those four categories — before asking the model to fix anything — is what makes the subsequent revision trustworthy rather than cosmetically improved.
For slide decks, the most common mistake is starting with design. A well-designed presentation with a broken argument is still a broken argument. The diagnostic move is to read the slide headlines in order, as a sequence of claims, before touching anything visual. If those headlines do not form a coherent case, better formatting will make the confusion look more professional. Structural revision — deciding what evidence matters, in what order — is genuinely hard reasoning, and worth spending on.
For documents, the most common mistake is asking for improvement before understanding the document’s actual job. Starting by asking the model to explain what the text requires the reader to do — not just what it says, but what it asks of someone who has to act on it — surfaces the editing work that actually needs doing. Qualifying an overconfident claim and moving the conclusion to the top are different jobs; “make this better” does not distinguish them.
The stale-data risk
One genuine risk in this pattern: low effort settings tend to skip search tools and answer from memory. For documents that depend on current filings, recent events, or live market data, that is a real problem — the draft looks complete while resting on outdated information.
The mitigation is a deliberate split: run any research turn at higher effort, then drop to low for file creation. Or instruct the model explicitly to search before building. Either way, the discipline is the same: treat effort as a setting you control per step, not a global dial you set once.
Where this leaves the practice
The broader shift the two-step pattern represents is a move from trusting the model’s depth to trusting the practitioner’s judgment about where depth is needed. The model can produce structure quickly; the human can identify, with some precision, which part of that structure is load-bearing and which is filler.
That allocation — cheap to build, targeted to sharpen — is likely to hold as models improve. The cost of low-effort output falls; the value of knowing exactly what question to ask with the expensive compute stays constant.
Text summarized and optimized using Anthropic’s models and reviewed by a human.