Near-Opus Pricing and the Tokenizer Tax
List price is a brochure. Dollars per accepted task is a business.
When a mid-flagship model ships with near-top quality and aggressive introductory rates, the feed fills with screenshots of the price table. That table is incomplete. Agent work is dominated by long traces, retries, tool results stuffed back into context, and tokenizers that may count the same English differently than the model you just left.
If you route half your desk to a "cheaper near-Opus" without measuring, you are not saving money. You are changing the unit of confusion.
The unit that matters
Stop optimizing for:
- dollars per million input tokens
- dollars per million output tokens
- "it feels cheaper"
Start optimizing for cost per accepted outcome: total dollars spent on a task class divided by how many results you would actually ship with light edit or less. Rejects still spent money. They belong in the real denominator even when they never appear in a vendor's happy dashboard.
A model that costs half per token and fails twice is not half price. It is a full-price failure with better marketing.
Where the tokenizer tax shows up
Tokenizers are not moral. They are accounting systems for text. A new tokenizer can make the same prompt more expensive even when the sticker rate falls. Community reports of higher burn on a new Sonnet-class generation are a hypothesis you should test on your traces — not a number you should paste as law.
How to test without mythology:
- Take ten real prompts from your last month (redact secrets).
- Run them on old default and new candidate with identical tools.
- Record tokens in/out, wall time, accept/reject.
- Compare cost per accept, not cost per call.
If you cannot get token counts, use dollars from the dashboard and the same ten tasks. Directionally honest beats precisely fake.
High-effort agent bills
Agent modes multiply everything. More reasoning tokens. More tool payloads. More "one more try." High-effort settings are a jet engine. Useful for architecture. Suicidal as a default for boilerplate.
Routing rule that keeps people solvent:
- Default effort for known patterns with tests
- High effort only when you would pay a senior human for the same hour
- Hard stop on steps, dollars, and wall clock before the agent starts
If you cannot state the stop conditions, you are not running an agent. You are running a tab.
The catch
Intro pricing expires. Habits do not. Teams that rewire defaults during a promo window wake up to a bill that matches the new normal, not the launch tweet.
The second catch is success theater. A long trace that eventually works still cost you the failures. Log rejects. They are the real unit economics of agents.
The third catch is comparing Opus-class and Sonnet-class on demo tasks that never hit tools. Tool-using work is a different sport. Benchmark the sport you play.
A one-page spreadsheet pattern
Columns:
- task_id
- model_id
- effort
- usd
- accepted (yes/light/no)
- notes
Weekly rollup:
- accepts
- total usd
- usd per accept
- reject rate
Promote a model only when usd per accept wins for two consecutive weekly boards on the same task mix. Demote without sentiment.
A note on "near Opus" marketing
Quality claims that sound like peerage ("near Opus", "almost frontier") are sales geometry. Sometimes they are directionally true on chat benches. Agent traces with tools are a different shape. Price the sport you play. If your desk is 80% tool-using coding agents, evaluate there. If it is 80% drafting, evaluate there. Do not let a general quality slogan set your default model for every lane.
Bottom line
Near-Opus is a quality claim. Intro pricing is a temporary coupon. The tokenizer is an accountant. Solo builders who fuse those three into one vibe will overpay.
Measure accepted outcomes. Cap effort. Re-run when the menu changes. That is how you keep a model ladder from becoming a subscription maze — and how you notice a tokenizer tax before the invoice becomes a personality trait.
Get new posts by email — first
The newsletter is in the works — join the waitlist and be first to know when it launches. Everything here stays free to read.
The Solo Stack is written by Matt — building products solo with AI, on his own infrastructure. If a claim isn’t backed by experience or a measurement, it doesn’t ship.
Not sending yet: joining stores your address on the waitlist. One confirmation email at launch — nothing sends unless you confirm.