Anthropic had told developers, in plain language, that Claude Sonnet 5’s introductory pricing would end on August 31. A 50% increase — from $2 to $3 per million input tokens, and from $10 to $15 per million output tokens — was scheduled to take effect the next day. Then, on August 10 and 11, the company reversed course. According to Anthropic’s own pricing documentation, that increase “will not occur.” The $2/$10 rate is now Sonnet 5’s permanent, standard price.
That reversal is the most consequential pricing decision to come out of the AI industry this month, and it landed in the middle of a broader stretch where nearly every major lab either cut prices or held them flat. OpenAI trimmed the cost of its flagship model. Google halved the price of its workhorse coding model. xAI’s Grok held steady. For a market that spent much of 2025 and early 2026 getting more expensive as capabilities climbed, August 2026 was, with one canceled exception, a month where AI got cheaper.
Table of Contents
What Happened
Four separate pricing and release decisions define the month, each made independently by a different company:
- Anthropic canceled a scheduled 50% price increase for Claude Sonnet 5, making its $2/$10 per-million-token introductory rate permanent (August 10–11).
- OpenAI cut the price of GPT-5.6 Sol, its flagship model, by 20% on input tokens and 33% on output tokens (August 21–22).
- Google launched Gemini 3.7 Flash at half the price of its predecessor, then retroactively applied that lower rate to the older model too (August 13).
- xAI — now operated under the SpaceXAI name following a February 2026 acquisition and July rebrand — released Grok 4.6 at unchanged pricing from the prior version (reported around August 12).
These aren’t coordinated moves — there’s no evidence any of the four companies were responding to each other’s specific announcements in real time. But taken together, they describe a market where competitive pressure is pushing prices down or keeping them flat almost everywhere, which makes Anthropic’s about-face particularly notable: it’s the one case where a price increase was actually on the calendar before being pulled.
| Model | Input (per 1M tokens) | Output (per 1M tokens) | Pricing Status |
|---|---|---|---|
| Claude Sonnet 5 | $2 | $10 | Permanent standard rate |
| GPT-5.6 Sol | $4 | $20 | Promotional through Nov 21, 2026 |
| Gemini 3.7 Flash | $0.75 | $3.75 | Introductory through Dec 31, 2026 |
| Grok 4.6 | $2 | $6 | Unchanged from Grok 4.5 |
Anthropic’s Reversal, in Detail
When Claude Sonnet 5 launched on June 30, 2026, Anthropic set its price at $2 per million input tokens and $10 per million output tokens, explicitly labeled as introductory pricing “through August 31.” The clear implication was that standard pricing — $3 and $15 respectively — would follow on September 1. Teams running Sonnet 5 in production had reason to budget for that jump.
They no longer need to. Anthropic’s official platform documentation now states that the introductory pricing is the standard price, and that the previously scheduled September 1 increase “will not occur.” For any business that had already built the higher rate into a Q3 or Q4 forecast, this is a straightforward, if quiet, cost win.
There’s a wrinkle worth understanding, though. Claude Sonnet 5 runs on a new tokenizer, and the same input text can now translate into more billable tokens than it would have under the previous Sonnet generation. The exact size of that effect isn’t pinned down by a single Anthropic-published figure — estimates from secondary sources vary — but the direction is consistent: some of the savings from the canceled price increase can be offset by simply needing more tokens to represent the same amount of text. In other words, the rate card didn’t change, but the bill for identical work still might, depending on what you’re sending the model.
OpenAI’s GPT-5.6 Sol Price Cut
OpenAI’s own developer documentation confirms that GPT-5.6 Sol, the top tier of its GPT-5.6 family, dropped from $5 to $4 per million input tokens — a 20% cut — and from $30 to $20 per million output tokens, a 33% cut. The change took effect around August 21–22, 2026.
Unlike Anthropic’s move, this one comes with an expiration date. OpenAI describes the new pricing as promotional, guaranteed to remain in place only through November 21, 2026. What happens after that hasn’t been announced. The GPT-5.6 family — Sol, along with the mid-tier Terra and budget-focused Luna — reached general availability on July 9, 2026, and had already seen a round of cuts to Terra and Luna in late July.
Google’s Gemini 3.7 Flash Launch
Google released Gemini 3.7 Flash on August 13, 2026, just three weeks after its predecessor, Gemini 3.6 Flash. The new model’s introductory pricing — $0.75 per million input tokens and $3.75 per million output tokens — is half of what 3.6 Flash originally cost. Google says that rate holds through the end of 2026, after which it doubles to $1.50 and $7.50.
In an unusual move, Google also applied the new, lower rate retroactively to Gemini 3.6 Flash, so the two workhorse models currently bill identically. Google positions 3.7 Flash for coding, agentic workflows, and document processing, and reports substantial internal benchmark gains over 3.6 Flash in those areas. Independent evaluation firm Artificial Analysis has also scored the newer model higher on its intelligence index and ranked it first among more than 180 models on output speed — a separate, independently sourced data point rather than a Google-reported figure.
Consumer access to Gemini 3.7 Flash runs through Gemini Spark, available to Google AI Pro and Ultra subscribers in more than 160 countries, though Google’s own rollout notes exclude the European Economic Area, the United Kingdom, Switzerland, and Nigeria at launch.
Other Notable Releases
Grok 4.6, from xAI — now operating as SpaceXAI after a February 2026 acquisition and rebrand — shipped at the same $2 input / $6 output per-million-token pricing as its predecessor, Grok 4.5. Coverage of the exact release date is inconsistent, ranging from August 6 to August 12; the later date appears better supported.
GLM-5.3, announced by Chinese AI lab Z.ai on August 14, is a more complicated case. The company says the model’s gains come entirely from additional post-training on the same base model used for GLM-5.2 — no new pretraining involved — and has published benchmark comparisons showing strong coding and cybersecurity results. Those figures are Z.ai’s own; no independent lab has audited them. More importantly, GLM-5.3 isn’t fully available yet — at launch, it’s accessible only through Z.ai’s subscription coding plans (starting at $18 a month) or its ZCode agent, with no public per-token API price and open weights pending a safety review expected roughly two weeks after the announcement.
What’s Still Unconfirmed
A few things worth flagging plainly rather than glossing over. Meta’s specific activity in the August 2026 pricing cycle is thin — one third-party tracker lists a model release attributed to Meta, but it hasn’t been corroborated through Meta’s own channels, so it isn’t treated as confirmed here. Reports of AI models exhibiting undocumented behavior changes or reduced effort — sometimes described informally as models getting “lazy” — circulate periodically online, but no such reports tied specifically to any of the August 2026 releases discussed above were found. And a claim that Anthropic is in acquisition talks with an AI infrastructure startup surfaced in a single source this month; without independent confirmation, it remains an unverified report rather than established fact.
What This Means for Your AI Budget
The net effect of August 2026 is a market that got cheaper or stayed flat almost everywhere a business might be spending. If you’d built a budget around a September 1 Sonnet 5 increase, you can revert that assumption. If you’re running GPT-5.6 Sol or Gemini Flash, your per-token cost already dropped — through late November and year-end, respectively.
The one caveat: a rate card isn’t the same as a bill. As the Sonnet 5 tokenizer change shows, actual token counts can shift independent of price. Check usage against the new rates rather than assuming the percentage savings apply one-to-one.
Frequently Asked Questions
Did the Claude Sonnet 5 price increase actually happen? No. Anthropic canceled the scheduled September 1 increase in early August. The $2 per million input token / $10 per million output token introductory rate is now Sonnet 5’s permanent standard price.
How much did OpenAI cut GPT-5.6 Sol pricing by? Input pricing dropped 20% (from $5 to $4 per million tokens) and output pricing dropped 33% (from $30 to $20 per million tokens), effective around August 21–22, 2026. OpenAI describes this as promotional pricing guaranteed through November 21, 2026.
Is Gemini 3.7 Flash cheaper than Gemini 3.6 Flash? Yes, at launch it was priced at half of 3.6 Flash’s original rate. Google then applied that same lower rate retroactively to 3.6 Flash, so the two models currently cost the same.
Why might my Claude Sonnet 5 bill still go up even though the price didn’t? Sonnet 5 uses a new tokenizer that can convert the same input text into more billable tokens than earlier Sonnet models did. The per-token price is unchanged, but the number of tokens you’re charged for a given amount of text may not be.
Is GLM-5.3 available through an API yet? Not as a standard, publicly priced per-token API at the time of writing. It’s accessible only through Z.ai’s subscription coding plans and its ZCode agent. The company has said open-weight and broader API access are pending a safety review.
Which frontier AI model is cheapest as of August 2026? This depends heavily on which pricing tier and provider you compare, and the landscape shifted multiple times within the month. Rather than naming a single “cheapest” model, the safest approach is to compare current published rates directly from each provider’s documentation at the time you need them.
When exactly did Grok 4.6 launch? Reporting is mixed. One tracking source places it on August 6, while more detailed independent coverage places it on August 12. The pricing itself — $2 input / $6 output per million tokens — is consistent across sources and unchanged from the prior version.

