Anthropic launched Claude Sonnet 5.5 on September 28, the second in the Claude 5.5 series, as the company prepares for an IPO. This marks an important shift in the criteria used to evaluate AI technologies: while raw benchmark performance remains a key consideration, companies now also care about how costly it can be to deliver results with a specific model.
Anthropic claims that Sonnet 5.5 operates faster by over 30% compared to Sonnet 5, and its application can cost up to 30% less for most tasks, even if token prices remain unchanged. According to Anthropic’s launch page, the said reduction in costs is achieved through reduced token usage for performing the same task.
Sonnet 5.5 has a pricing of $2 per million input tokens, $10 for output tokens used, and $0.20 for cache-read tokens used. Opus 5.5, which specializes in complex reasoning and judgment operations, is priced at $4 and $20 for input and output tokens, respectively. Sonnet is meant for simple tasks like coding, document generation and spreadsheets.
This raises efficiency to the forefront of the conversation. For many tasks, businesses do not have to go for the highest-performing model. If a low-cost model can complete the desired task consistently, that can be more significant than scoring a few points higher in a benchmark. This aspect is especially important for Anthropic, as Reuters reports that around 80% of the company’s revenue comes from enterprise customers.
Anthropic says Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, compared with 10.3% for Sonnet 5. But the Artificial Analysis leaderboard shows why benchmark scores are only part of the picture.
Sonnet 5.5, at maximum effort, reaches a score of 56 on the intelligence scale, costing an estimated $7.60 for each individual task. GPT-6 Sol can achieve 48 points at $1.06 per task, while Xiaomi’s MiMo-V2.6-Pro achieves a score of 46 for a paltry cost of only $0.13. Such differences in cost motivate companies to favor multiple routes: one can assign harder tasks to premium-grade systems while all routine tasks can be done by cheaper systems.
This change falls into a much larger trend seen over time. According to Epoch AI’s report from September 22, the costs for a specific level of AI performance have decreased roughly 47% each quarter starting in 2023, an approximately 13-fold yearly reduction.
A Cryptopolitan report published on September 29 states that corporate buyers are increasingly favoring the most affordable model that accomplishes the required tasks, while low prices also allow businesses to justify the use of multiple models.
However, lower costs do not ensure lower AI expenses. Gartner’s report from August 17 titled “Gartner Predicts AI Inference Costs Per Agentic Workflow Will Increase More Than Fivefold Through 2028” talks about the “inference paradox.” The decrease in the price of tokens stimulates more complex workflows, which results in an increase in total token usage.
“Product leaders cannot rely on more efficient token economics to rationalize AI costs.”
— Will Sommer, Senior Director Analyst at Gartner, in Gartner’s August 17, 2026 report,
According to the Global AI Diffusion Report published by Microsoft, low-cost and open-weight models have the potential to increase accessibility, particularly in the Global South. While it is true that lower model costs will improve accessibility, infrastructure, connectivity, and skills are deciding factors for whether cheaper AI models will be used.
For enterprise buyers, Sonnet 5.5 shows where the market is heading: the best model is not necessarily the most powerful one, but the one that can do the job well at a reasonable cost. Haiku 5.5, due in the coming weeks, could push that shift even further by making high-volume AI work cheaper to perform and harder for businesses to justify paying higher prices for routine tasks.
The smartest crypto minds already read our newsletter. Want in? Join them.