AI Just Got 50x Cheaper This Week. The Skill Gap Just Got Wider.
DeepSeek's new model costs a fraction of a cent per million tokens, and OpenAI just cut its own prices 80% in three weeks. The tasks got cheaper — knowing which tier each one needs didn't.
By Patin Team · Examples are illustrative composites
A capable AI model now costs less than a third of a cent per million tokens. If your team is still routing every task through the most expensive model available, you're not being careful — you're paying frontier prices for judgment calls a cheaper model would get right just as often.
What shipped and why it matters
DeepSeek released V4-Flash on July 31: a 284-billion-parameter model that activates only 13 billion of them per request, scores 82.7 on Terminal-Bench 2.1, and costs $0.28 per million output tokens — roughly 10 to 50 times cheaper than a frontier model for routine agent work. Three days earlier, OpenAI cut GPT-5.6 Luna's input price 80%, to $0.20 per million tokens, less than three weeks after launch. TLDR AI reported the same week that open-weight models like GLM 5.2 and Kimi K3 now match proprietary models on regulatory tasks at roughly a third of the cost.
Capability isn't the bottleneck anymore, and access was never really the bottleneck either — a Stanford SIEPR brief published July 26 found that firms adopting AI expanded headcount by 10%, not shrank it. What's left is knowing which model tier a given task actually needs.
The skill that's left
That's a specific, learnable skill, and it isn't one the pricing page teaches you. The default habit is whichever model the company licenses, or whichever one worked last time, regardless of what the task calls for. A prompt that reformats a spreadsheet doesn't need the same reasoning depth as one drafting a client-facing risk memo — but without separating those two categories, you're either overpaying on the easy ones or under-serving the hard ones.
Monday morning, the fix is a five-minute audit: pull your last ten AI tasks and sort them into three piles — routine, analytical, and high-stakes — then check whether the model you used actually matched the pile.
Two ways this goes wrong
Dana runs marketing operations at a 40-person SaaS company. Her team burns most of its AI budget on the flagship model for tasks that don't need it: reformatting campaign briefs, pulling summary stats from weekly reports, drafting first-pass social captions. None of that requires frontier-level reasoning. Routed through a model like DeepSeek V4-Flash or GPT-5.6 Luna instead, those tasks cost a fraction of what they did last quarter — freeing budget for the handful that do need the expensive model, like competitive positioning memos or messaging for a product pivot, where a wrong call is expensive to unwind.
Marcus, a solo analyst at a regional bank, hit the opposite problem. He'd been running every client-facing summary through the cheapest available model to keep costs down, including drafts that fed into compliance reports. Cheap model, careless output — a fabricated figure made it into a draft before a colleague caught it during review. The lesson isn't "always use the expensive model" any more than it's "always use the cheap one." Reasoning depth and cost are separate dials, and matching them to what a wrong answer would actually cost is the skill the last month of price cuts just made non-optional.
The tools got cheaper across the board this week. The judgment about which one to use for which job is the only cost that didn't drop — and it's the one most people haven't started paying attention to.
Put this into practice
Reading is a start — but skill comes from doing. Try these drills now.
Reading about it only gets you so far
Patin turns this into five-minute drills that score what you write and tell you why. It's in closed beta — join the waitlist and we'll email you when your cohort opens.
Just want the writing? .
Keep reading on this
Claude Now Has a 'Think Harder' Button. Here's When to Press It — and When Not To.
Anthropic shipped an effort toggle with Claude Opus 5. Google released three tiers of Gemini Flash and deprecated temperature settings. Every major AI provider made the same call this week: match the task to the reasoning depth. Here's how.
5 min readYour AI Just Got an Effort Dial. Here's When to Turn It Up.
Claude Opus 4.8 shipped with explicit effort controls — users choose whether the model reasons carefully or answers fast, at 3x the cost difference. Here's how to decide which tasks deserve which mode.
4 min readYour AI Vendor Just Drew an Ethical Line. Here's Why That Affects Your Workflow.
Court documents unsealed July 2 showed Anthropic drew two non-negotiable redlines — no mass surveillance, no autonomous weapons — and lost government access because of it. The 19-day Fable 5 suspension was the operational consequence. Here's what it means for how you build with AI.
5 min read