Which Claude AI Model to Use (Stop Wasting Tokens)

· Claude Code

The fastest way to answer which Claude model to use is to stop reading benchmark charts and start thinking like a hiring manager with a payroll to protect.

Which Claude Model to Use, Framed as a Payroll Decision

Most model comparison articles give you the same thing. A spec table, a benchmark score, a shrug. You learn what the models are without learning anything about your own work, so you go back to running everything on the most expensive one because it feels safer. That instinct is what drains your usage limits by Tuesday afternoon. Here is the reframe that fixed it for me. Your models are not tiers on a pricing page. They are people on a team, and you are the one signing the checks. Once you see the lineup that way, the routing decision answers itself. Nobody pays a doctor to organize a filing cabinet. Yet that is exactly what you do every time you ask a frontier model to reformat a CSV.

Fable is the founder

Fable 5 sets direction. It is Anthropic's most capable widely released model according to the official model overview, and you bring it the hard architectural call, the messy strategic decision, the problem where you genuinely do not know the shape of the answer yet. It scopes the work and hands off the rest. You do not ask your founder to write the weekly recap email. You ask them what the company should be doing next quarter.

Opus is your $500-an-hour expert

Opus 5 is brilliant, costly, and available in limited hours. It is the specialist you bring in for the thing that actually matters, then release before the meter runs your budget dry. Real client work, a landing page that has to convert, copy a customer will actually read. That is Opus territory. Not your morning summary. If you are a marketer rather than an engineer, this is the tier that earns its keep on positioning, offer copy, and campaign strategy. The rest of the week belongs further down the ladder, which is the practical half of why marketers should use Claude Code in the first place.

Sonnet is the solid $50-an-hour employee

Sonnet 5 is good at most things. It handles the broad middle of the job: drafting, working through problems in steps, writing code, the daily volume that makes up most real work. This is the hire that quietly does the bulk of your output. Most people underuse it because it lacks the prestige of the top tier.

Haiku is the $10-an-hour assistant

Haiku 4.5 is fast and excellent at simple checking. Classification, extraction, formatting, routing, is-this-thing-what-I-think-it-is. Anthropic's own guidance on choosing a model recommends Haiku for simple tasks and reserves the top tier for the most complex reasoning. Speed matters more than depth here, and Haiku has speed.

Codex is the contractor from a different agency

Codex, running GPT-5.6, sits outside your org chart entirely. That is the whole point. Hiring an outside contractor does not burn your own team's hours. I use it as a second set of eyes on writing. It reads a draft, tells me where the argument sags, and none of that comes out of my Claude budget. Different agency, separate invoice.

What the Cost Ladder Actually Looks Like

The analogy earns its keep only if the numbers back it. They do, though not in the order most people assume. Here is what Anthropic charges per million tokens, straight from the official Claude API pricing page:
ModelInputOutputThe role it plays
Claude Fable 5$10$50Founder
Claude Opus 5$5$25Senior expert
Claude Sonnet 5$2$10Core employee
Claude Haiku 4.5$1$5Assistant

 

Claude model pricing compared: Fable 5, Opus 5, Sonnet 5, and Haiku 4.5 input and output cost per million tokens
The same four models as a payroll. Fable sits at the top at $10 in and $50 out per million tokens, ten times what Haiku costs.
One correction worth making, because most people have this backwards. Fable is the priciest widely available seat in the building, at twice the cost of Opus. Not the other way around. Opus is the second-most expensive model on the list, not the top of it. The spread that should actually change your behavior is Fable to Haiku. That is 10x on input and 10x on output. Run a batch of 500 classification jobs on Fable instead of Haiku and you paid a founder's rate to sort mail. Note that Sonnet 5 is at introductory pricing of $2 and $10 through August 31, 2026, then moves to $3 and $15. That narrows the Sonnet-to-Opus gap from 2.5x to closer to 1.7x, so check the date before you treat the table as permanent. One more distinction the pricing page will not make for you. These are API rates, billed per token. If you work inside a Claude subscription plan instead, you are spending against usage limits rather than a card, and the ratios still hold even though the units differ. Either way the expensive model consumes your capacity faster.

How I Actually Route Work Across Claude Models

A framework is easy to write and harder to run. Here is the routing I use in a normal week:

 

That is the whole rule. The rest is knowing which bucket a task belongs in.
Claude model routing workflow: plan with Fable or Opus, build with Sonnet, check with Haiku, then verify with Opus
The shape that matters. Cost dips through the middle of the job and climbs back up for verification, because checking work is judgment and not clerical.
Planning and hard calls go to the top tier. When I am scoping a new system, deciding how a workflow should be structured, or working through something I have not solved before, I want maximum judgment. That is a small share of my total spend and where most of the payoff shows up. The work itself drops down a tier or two. Once the plan exists, the work is mostly mechanical. Drafting, editing, running the steps. Sonnet handles that volume at a fifth of the top-tier rate, and the quality gap on well-scoped work is much smaller than the price gap suggests. Verification climbs back up. This is the part people miss. Checking work is a judgment task, not a clerical one. I plan high, run the work cheap, and verify high again. If you do the whole loop on one model you either overpay for the middle or underthink the ends. Simple checks stay at the bottom. Does this file have the field I expect? Is this post over the character limit? Did the scan return clean? Haiku answers those in a second for pennies, and sending them to a top-tier model buys you nothing but a bigger bill. This is also the logic behind running Claude Code agent teams, where cheap sub-agents do the reading and the expensive one keeps the judgment. The common failure is defaulting upward on everything. It feels like insurance. It is really just paying a specialist to do intake paperwork, and you only notice once the capacity is gone. If you want the tactical version of this, I wrote up ten Claude Code usage tips that cover the token math in more detail.

Where the Hiring Analogy Breaks Down

I would rather hand you the limits than pretend the frame is perfect. An employee learns. Models do not, at least not between sessions. Your $50-an-hour hire gets better at your business over six months. Sonnet starts fresh every time unless you give it the context yourself, which is why your prompts and files do the work a real onboarding would. Cheaper models are not simply worse. Anthropic positions Haiku 4.5 as near-frontier performance at its most economical price point, built for high-volume straightforward work, and the Haiku 4.5 release put it within reach of Sonnet 4 on coding at a third of the cost and more than twice the speed. It is not a junior version of the same brain. On the narrow work it was built for, the tradeoff you think you are making mostly does not show up. The pricing table also reflects raw token cost only. Fewer, better attempts from a stronger model can beat many cheap retries, so a model that gets it right on the first pass sometimes costs less than one you correct three times. I hit that line myself when Claude Fable 5 launched, and my honest read was that it was overkill for 90 to 95% of my marketing and content work.

Ryan's Final Thoughts

Model selection is a management problem wearing a technical costume. Here is the part most people never accept: the price ladder is not a quality ladder. Paying more does not buy you a better answer on a task that was never hard. It buys you a slower, costlier version of the same answer, which means defaulting upward is superstition rather than insurance. Route by what the task actually demands. Plan at the top, work in the middle, check at the bottom, and send the overflow to a different agency entirely.

Which Claude Model to Use FAQs

Which Claude model should I use for most tasks?

Sonnet 5 for the majority of daily work. It handles drafting, multi-step reasoning, and code generation at $2 per million input tokens and $10 per million output, which is a fifth of Fable's rate. Move up to Opus 5 for work a client will see, and up to Fable 5 for genuinely hard architectural or strategic calls.

Is Claude Opus worth the extra cost over Sonnet?

For a specific slice of work, yes. Opus 5 runs $5 input and $25 output versus Sonnet's $2 and $10, so you pay roughly 2.5x. That is worth it for novel problems, decisions with many competing constraints, and final quality checks. It is not worth it for drafting, summarizing, or formatting.

What is Claude Haiku actually good for?

Classification, extraction, formatting, routing, and simple verification. At $1 input and $5 output it is the cheapest model in the lineup, and Anthropic builds it for high-volume straightforward tasks where speed and price matter most. Use it for mechanical checks that do not need deep reasoning.

Should I use GPT alongside Claude models?

It helps for one practical reason. Running a GPT model through Codex bills against a separate account, so a second-opinion pass does not eat into your Claude capacity. I use it to critique drafts and catch weak arguments, then decide myself what to accept. Its read is input, not a verdict.

Does using a cheaper Claude model hurt output quality?

Not much, on tasks that fit the model. The quality gap tends to show up on open-ended reasoning and matters far less on well-scoped mechanical work. The bigger risk runs the other direction. Paying top-tier rates for routine jobs burns your capacity so fast that you have nothing left for the tasks that genuinely need judgment.

40+ Claude Code skills that run my business. One purchase, lifetime updates.

Get instant access

Free AI Marketing Guide

Get the free AI Marketing Guide

The exact AI tools and Claude Code skills I use to run my agency solo.

No spam. Unsubscribe anytime. Privacy policy.