Three frontier models shipped inside 48 hours, and two of them cut prices by half. Here's what each one actually costs to run an agent, why the model with the cheapest output tokens isn't the cheapest agent, and where GPT-6 Luna at ten cents fits in.
Anthropic shipped Opus 5.5 on 22 September. About ninety minutes later, OpenAI shipped GPT-6 Sol and Luna at half their previous prices. Grok 4.7 had landed the day before.
If you run agents, your model budget changed three times in two days and nobody sent you a memo.
So here's the memo. Every price below is from the vendors (our LLM pricing guide has the wider grid), checked 28 September 2026, and the per-task maths is worked out so you can redo it with your own numbers.
What each model costs
| Per million tokens | Opus 5.5 | GPT-6 Sol | Grok 4.7 | GPT-6 Luna |
|---|---|---|---|---|
| Input | $4 | $2 | $2 | $0.10 |
| Output | $20 | $10 | $6 | $0.50 |
| Cache reads | $0.20 | $0.20 | $0.50 | $0.01 |
| Context window | see note | 1.05M | 500K | 1.05M |
| Released | 22 Sep | 22 Sep | 21 Sep | 22 Sep |
Grok 4.7's rates apply below 200,000 prompt tokens. Cross that line and every rate doubles, on the whole request. Opus 5.5's context window wasn't stated in the launch coverage we could verify, so we've left it out rather than guess.
Three things behind that table are worth knowing.
Opus 5.5 cut cache reads by 60%, from $0.50 to $0.20, which matters far more than the 20% sticker cut because cache reads dominate agent spend. OpenAI went further and gave cached input a 90% discount across GPT-6, and says cache hits now survive a change in reasoning effort or a tool being switched on, which used to blow the cache away. And OpenAI told VentureBeat the new rates are permanent, not promotional.
Grok is the odd one out. Its output is the cheapest of the three frontier models at $6, but its cache reads are $0.50, two and a half times what Opus 5.5 and Sol charge. Hold that thought.

Cost per agent task
Take a normal agent run: about 40,000 tokens of cached context (instructions, tool schemas, history), 3,000 tokens of fresh input, 2,000 tokens of output.
| Opus 5.5 | GPT-6 Sol | Grok 4.7 | GPT-6 Luna | |
|---|---|---|---|---|
| Cache read (40K) | $0.008 | $0.008 | $0.020 | $0.0004 |
| Fresh input (3K) | $0.012 | $0.006 | $0.006 | $0.0003 |
| Output (2K) | $0.040 | $0.020 | $0.012 | $0.0010 |
| Per task | $0.060 | $0.034 | $0.038 | $0.0017 |
| 200 runs/day | $360/mo | $204/mo | $228/mo | $10/mo |
And there's the thing nobody's saying out loud.
Grok 4.7 has the cheapest output tokens of the three frontier models and still costs more per agent task than GPT-6 Sol. Its cache reads are 2.5 times higher, and on agent work the cache line is most of the bill.
That's the whole argument for reading the cache row first. On a chat product, where each message is mostly fresh input, Grok's cheap output wins. On an agent, where you re-send the same 40,000-token prefix on every call, it loses to a model with pricier output and cheaper cache.
One caveat that goes the other way: route Grok 4.7 through OpenRouter and it lists at $1.60 input, $4.80 output and $0.40 cached, below xAI's own card. At those rates the same task costs about $0.030, which puts Grok back in front. Worth checking before you conclude anything from the list prices alone.
GPT-6 Luna at ten cents: can it actually run an agent?
Luna is the interesting one, and the number looks unreal next to the others. $0.10 input, $0.50 output, one cent per million on cached reads, with the same 1.05 million token context window as Sol.
On our example task that's $0.0017 a run. Twenty times cheaper than Sol. Thirty-five times cheaper than Opus 5.5. You could run that agent 6,000 times a month for about ten dollars.
So what's the catch? Luna is a budget tier, not a frontier model in disguise. It's the model OpenAI is making the default for ChatGPT Free and Go users, which tells you where it sits. For agent work that means it's a fine fit for the jobs that make up most of an agent's day, and a bad fit for the ones you remember:
- Good for Luna: classification, labelling, triage, extraction, short summaries, routing decisions, simple tool calls with a clean schema.
- Not for Luna: long multi-step chains, code that has to compile, anything where a wrong answer is expensive to unwind.
The honest version is that we can't tell you Luna's tool-calling reliability from a spec sheet, and neither can anyone else who hasn't run it. Fifty identical tool calls against a strict schema is a twenty-minute test, and it's the test worth running before you move a production agent. Until someone publishes that, treat the price as an invitation to try it on your cheapest workload, not as a reason to migrate.
At $0.0017 a task, Luna changes what's worth automating. Jobs you'd never write an agent for because the tokens cost more than the time saved are suddenly worth a try.
So which one should your agent use?
Not one of them. That's the actual answer, and it's the part most comparisons skip.
GPT-6 Sol is the best default for agent work right now. It's the cheapest of the three frontier models per task, it has a 1.05M context window, and its caching behaviour is the friendliest to long-running agents. OpenAI claims Sol beats Claude Opus 5 on AutomationBench at roughly 9% of the cost per completed task, which is a vendor benchmark, but the price maths above is independent of it.
Opus 5.5 earns its 1.8x premium on long, multi-step coding work where finishing correctly matters more than the token bill. Anthropic reports it at default effort beating Opus 5 at max effort on Terminal-Bench 4.0 for about a fifth of the cost. One thing to know before you point an agent at it: Opus 5.5 ships with classifiers that can hand a flagged request to an older model without telling you in the response body, which we covered in the Opus 5.5 rerouting problem.
Grok 4.7 makes sense if your prompts are short and output-heavy, which is the opposite of the classic agent shape, or if you're buying it through a router at the lower listed rates. Watch the 200K cliff: cross it and the whole request bills at double.
Luna belongs in the routing table, not on its own. Put the fifty-a-day jobs on it and keep a frontier model for the rest.
Which is really the point. The gap between $0.0017 and $0.060 a task is 35x, and no single model is right for both ends of that range. A router that sends classification to Luna and refactors to Opus 5.5 pays for itself in a day. We wrote the logic up in the model routing setup guide, and if cache reads now dominate your bill, prompt caching for agents is where the rest of the money is hiding.
On BetterClaw, the model is a per-agent setting and you bring your own keys, so switching an agent from Sol to Luna is one dropdown and the provider bills you directly with nothing added on top. Plans start at $19 a month.

What two days of price cuts actually mean
Six months ago, running a frontier model on every agent task was a real budget decision. This week Sol landed at exactly Sonnet 5's price, Opus got 40% cheaper to run, and OpenAI put a usable model on the table at a tenth of a cent per thousand tokens.
The cost of being wrong about your model choice keeps falling. The cost of not knowing which model your agent is using, or what it bills you per task, stays exactly the same.
So don't migrate everything this week. Work out what one task costs you today, then check it again after you switch. That number is the only one that survives the next launch, and on current form the next launch is about eleven days away.
If you'd rather change models with a dropdown than a deploy, that's what we built. Every agent picks its own model, you bring your own API keys with no markup, and per-agent spend caps mean a routing mistake costs dollars instead of a weekend. Plans start at $19 a month for Basic, $49 for Pro and $149 for Business. Start now or see the full pricing.
Frequently Asked Questions
What is the cheapest AI model for agents in 2026?
For frontier-class work, GPT-6 Sol at $2 input and $10 output per million, which works out around $0.034 per agent task on a typical run. For high-volume simple work, GPT-6 Luna at $0.10 and $0.50 is dramatically cheaper at roughly $0.0017 a task. Most teams end up using both, routing simple jobs to Luna and keeping a frontier model for complex ones.
How does Opus 5.5 compare to GPT-6 Sol for agent work?
Sol is cheaper: about $0.034 per task against $0.060, with the same $0.20 cache reads and a 1.05M context window. Opus 5.5 justifies its premium on long multi-step coding, where Anthropic reports it beating Opus 5 at max effort for about a fifth of the cost. Sol is the better default; Opus 5.5 is the better escalation.
How much does GPT-6 Sol cost compared to Grok 4.7?
Sol is $2 input and $10 output per million; Grok 4.7 is $2 and $6 on prompts under 200,000 tokens. Grok looks cheaper until you count cache reads, where it charges $0.50 per million against Sol's $0.20. On a cache-heavy agent task Sol comes out ahead at $0.034 versus $0.038, and Grok's rates double on the whole request once a prompt hits 200,000 tokens.
Is GPT-6 Luna good enough for agent tool calling?
For simple, well-structured tool calls, probably. For long chains, treat it as unproven. Luna is OpenAI's budget tier and is becoming the default for ChatGPT Free users, so it isn't a frontier model at a discount. Run fifty identical tool calls against your real schema before trusting it with anything unattended; that test takes twenty minutes and tells you more than any benchmark table.
Is it worth switching models every time prices drop?
No. Switching has a real cost in prompt tuning, eval runs and reliability surprises, and three models repriced in two days. Work out your cost per completed task rather than per token, switch when the gap is large enough to matter, and keep the ability to change models per agent so the next launch is a config change rather than a migration.




