For a long time, the conversation about which AI model to use was almost entirely about benchmarks. Which one scored highest on reasoning tests? Which one passed the bar exam? Those numbers matter — they’re not noise — but they answer a different question than the one most practitioners are actually asking.
The question we’re actually asking is: which model will make my work go better?
The gap between “smarter” and “better for you”
Running a benchmark suite across the models in regular use reveals a consistent pattern. By the averages, Fable 5 — Anthropic’s latest model — clears eighty or above across every tracked dimension. GPT-5.6 Sol (OpenAI’s current flagship for knowledge work) scores higher in some areas and lower in others, with a particular dip on anything involving messy, unstructured data.
And yet Sol is still the model opened every morning.
That isn’t stubbornness or brand loyalty. It’s something more specific: the way the work is structured happens to match what Sol does well. My prompts tend to be long and front-loaded. The edges get stated upfront. When a draft is inspected and something is off, the expectation is that the model changes exactly that thing and leaves the rest alone — literal obedience, not a reinterpretation. Sol is very good at this. It takes a detailed brief and executes it faithfully. For my workflow, that matters more than the aggregate score.
Briefers and finders
Here’s a distinction that’s been useful to me. Some models behave as briefers — they take instructions at face value, follow through precisely, and respond to corrections literally. Others are more like finders — they read between the lines, infer what you probably meant, and surface an angle before committing to a draft.
Fable sits firmly in the finder camp. When the important work is still happening between your intent and your words — when you have a direction but not yet a story — Fable is genuinely useful in a way Sol isn’t. It doesn’t wait for you to fully specify the brief; it helps you find the brief.
Neither disposition is better. They’re suited to different moments in the work, and to different people doing ostensibly the same job.
Same task, different models
Consider two people both building a launch plan from the same customer research. One of them knows the narrative — they’ve read the research, they know which problem leads, they just need execution. The other is still working out which angle is worth leading with.
Both people are doing “product marketing.” A task-based routing system would send them to the same model. But the hard part of the work is different for each of them, and the model that helps one person will frustrate the other.
This is why “start with the job” — the common advice for model selection — is only half the answer. The other half is the person doing the job, and specifically, the shape of how they work.
Working backwards from what went well
The most useful diagnostic available is this: think of a piece of work that went unusually well with AI assistance, and walk it backwards. Did you know what you wanted to say at the start, or did you discover it through the drafts? Did you brief fully upfront, or did you react to what the model produced? When something was wrong, did you want the model to fix exactly that thing, or did you want it to understand what you meant?
Those answers tell you more about your working style than any personality quiz. They also tell you which model family is likely to be your daily driver.
The rotation underneath the driver
A daily driver isn’t the whole story. Even if Fable or Sol is where you start each morning, there are situations where neither is the right call. For anything where access to live, current information is the whole point of the task, a search-first model matters more than reasoning depth. For high-volume, clearly-specified work where the judgment is already done and you just need throughput, a faster and cheaper model starts to make sense. Reaching for a different model in those moments isn’t infidelity to your workflow — it’s just using the right tool.
The rotation looks something like: one daily driver that fits your working style, one backup with a specific trigger situation where it outperforms your default, and a handful of situational tools for specialist needs.
Where this goes from here
The better the models get, the less the aggregate benchmark will distinguish them for most people — and the more individual working-style differences will drive which one actually helps. We’re already partway there. The interesting question is no longer “which model is smartest” but “which model fits the way we think.”
That’s a question worth spending some time on. The answer is more stable than the model release cycle.
Text summarized and optimized using Anthropic’s models and reviewed by a human.