For years, AI tools were approached the same way most people do — as a Swiss Army knife you reach for when you need to move faster. The question was always “which tool is smartest?” A more useful question is: “what kind of problem is actually being solved?” That shift changes how to think about the work, and where AI fits into it.
A useful taxonomy has emerged from watching how the current generation of models is actually being deployed: six axes of difficulty that capture why knowledge work is hard. Not all six are moving at the same speed. Understanding which ones are automating now — and which ones are not — is more practically useful than any benchmark comparison.
The six reasons things are hard
Most knowledge work is hard for one or more of these reasons:
Reasoning is multi-step logical deduction from first principles. Think: structural tax analysis, complex regulatory interpretation, building a pricing model for a novel financial instrument. The thinking on each step is genuinely difficult.
Effort is not intellectually hard, just enormous in scale. Reviewing 3,000 contracts. Auditing every customer interaction from last quarter. The challenge is endurance across a huge surface area, not the quality of thought on any single step.
Coordination is routing work, aligning teams, managing information flow so the right people know the right things at the right moment.
Domain expertise is pattern recognition built from lived experience. The senior engineer who debugs faster because she has seen that error before. The attorney who reads a deal better because she has closed 300 of them. The gap between “has read about it” and “has done it” is still real.
Emotional intelligence is reading the room, delivering difficult feedback, navigating relationships where the stakes are personal. No model handles this reliably.
Ambiguity and judgment is figuring out what the question even is. Deciding what to build when the market signal is contradictory. Choosing the strategically right but politically uncomfortable path.
Which ones are moving this quarter
The axes are not all moving at the same speed, and that gap is more clarifying than most things written about AI this year.
Reasoning is getting automated fast. Recent benchmark improvements on novel logical reasoning — the kind that requires working from first principles rather than pattern-matching against training data — have shown gains in months that used to take years. The practical implication: tasks that are primarily reasoning-bottlenecked are starting to yield, especially in mathematical, scientific, and highly structured domains.
Effort is already being automated at scale. Agentic pipelines — AI systems that can sustain work across many hours and thousands of items — have made the “large but not intellectually complex” category genuinely addressable right now. If most of your value sits on this axis, the timeline for disruption is not three to five years. It is now.
Coordination is starting to move. The ability of AI systems to use tools, track state, and route work across organizational context is improving, though it still requires significant scaffolding to work reliably in real business environments.
The other three — domain expertise, emotional intelligence, and ambiguity and judgment — are, in my observation, barely touched. There are good structural reasons to think they will be the last to yield, if they yield at all.
Why this matters for how we think about our own roles
Running through a typical work week with this taxonomy in mind, the first thing that stands out is that the axis mix is rarely what most people assume. Most roles spend far more time on coordination, emotional intelligence, and judgment than on pure reasoning — even technically sophisticated roles. A tax attorney who seems like a pure reasoning job might spend only ten percent of her week on genuine multi-jurisdiction analysis. The rest is client management, document gathering, and judgment calls in ambiguous situations.
This matters because without this map, it is very hard to make useful predictions about which parts of your value are durable and which parts are dissolving.
The skill that compounds across all six axes
The question worth returning to is what might be called the taste question: as AI systems generate better-looking output across all six axes, the value of being able to evaluate whether that output is actually good increases, not decreases. A model that produces a confident but subtly wrong financial model is more dangerous than no model at all, if no one on the team can spot the flaw.
Model fluency matters. But the skill that will compound over the next few years is the judgment to evaluate AI output in a specific domain. Not prompt engineering as a craft in itself, but the domain expertise and critical instinct to know when a plausible-looking answer is actually wrong.
That distinction, between generating and evaluating, is where the real professional leverage sits for some time.
Text summarized and optimized using Anthropic’s models and reviewed by a human.