The Best AI Model for Your Work Isn't the Best for Your Thinking
Ethan Mollick's research found the models best at doing your work alone score lower on helping you think it through, while a cheaper model does both well. Here's the three-question check for matching model to task.
By Patin Team · Examples are illustrative composites
If you default to the most expensive AI model for everything because it's "the best one," you're optimizing for the wrong thing on a chunk of your tasks — the model that works best alone isn't the model that helps you think best.
Two different jobs, two different winners
Ethan Mollick shared research this week (Aug 25) comparing how well models perform two distinct jobs: doing a task autonomously, and helping a person do that task better. Claude's Opus and Sonnet models scored well on the first — hand them a task and they execute it competently with no input from you. But they scored lower on the second: augmenting a human's own work, catching what you missed, pushing back on a weak plan. GPT-5-Mini, a smaller and cheaper model, performed well on both.
That split matters because the default habit is to open whatever tool is on screen and use it the same way for drafting an email and for stress-testing a strategy, as if "smart" were one dial instead of two different skills.
The same week, a Hacker News post titled "Small Models Have Arrived" (784 points) argued that roughly 60% of routine business tasks can now run on models costing about $0.10 per query, versus $1 or more for frontier models — with comparable results on tasks that don't need deep reasoning. And Simon Willison's write-up of the Ramp AI Index (Aug 23) showed enterprise buyers already acting on this: Anthropic's Opus 4.8 holds 28% share of enterprise usage against 8% for the pricier Fable 5. Companies spending six figures a year on AI are already routing routine work to the cheaper option.
The three questions
Before you open your AI tool, ask what you're actually hiring it to do:
- Do I want it to do this for me, or help me do it better? A first draft of a routine email is a "do it" task. Pressure-testing next quarter's plan is a "help me think" task. They can call for different models.
- Does this need a frontier model, or would a cheaper one produce the same result? If the "small models" estimate is even close to right, most of what lands in your AI tool this week is routine enough to run on the cheap tier — summarizing, formatting, first-pass drafting. Save the expensive model for the tasks where reasoning depth actually changes the output.
- What am I optimizing for — speed, cost, or collaboration? If the honest answer is "I just want this done," a cheap model that executes reliably beats an expensive one that argues with you.
Two teams, two defaults
A communications lead at a 25-person nonprofit ran every task through the same premium model subscription: press releases, donor emails, and also the annual strategy memo she was drafting for the board. The press releases came out fine — routine writing, low stakes if a sentence needed a tweak. The strategy memo didn't: the model produced a confident, well-organized document that never questioned her assumption about which funder segment to prioritize, because that wasn't the job it was good at. She caught the gap only when a colleague asked a question the AI never had.
A finance manager at a 90-person logistics company took the opposite approach. Routine work — expense categorization, monthly variance summaries — went to the cheapest model that could handle structured, repetitive tasks, at a fraction of the per-query cost. The one task he routed to a frontier model, deliberately, was reviewing a vendor contract renewal where he wanted the AI to argue back if his read of the terms looked wrong. Same team, same total spend on tools; the difference was matching the model to which of the two jobs — doing or helping — each task actually needed.
The takeaway
"Best model" isn't one ranking — it's at least two, and they don't always agree. The skill worth building isn't picking the most capable tool available; it's noticing which job you're actually asking AI to do before you pick a model to do it.
Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
The Model You Pick Matters Less Than the Job You Give It
Four separate AI launches this year taught the same lesson from different angles: model choice matters less than task fit. Here are the three questions to ask before you pick a model for anything that matters.
4 min readEvery AI Tool Now Asks How Hard to Think. Most People Don't Answer.
Claude, Gemini, and DeepSeek all shipped the same change this year — explicit control over how hard the model thinks, priced across a 50x range. Here's the three-tier rule for matching effort to what a task actually needs.
4 min readAI Just Got 50x Cheaper This Week. The Skill Gap Just Got Wider.
DeepSeek's new model costs a fraction of a cent per million tokens, and OpenAI just cut its own prices 80% in three weeks. The tasks got cheaper — knowing which tier each one needs didn't.
4 min read