The AI Race for Cost and Autonomy: OpenAI GPT-6, Claude Opus 5.5 and Meta Muse
The AI race for cost and autonomy is entering a new phase. OpenAI is launching cheaper GPT-6 Sol and Luna models, Anthropic is answering with Claude Opus 5.5, while Meta is discovering that even an advanced AI agent can still need a human being. The market is changing: building the smartest artificial intelligence is no longer enough — it also has to be fast, affordable and capable of completing real-world tasks on its own.
Until recently, the biggest AI companies competed mainly on benchmarks: who could solve a harder problem, write better code or score a few points higher on a test.
Now economics is moving to the center of the race.
OpenAI and Anthropic are almost simultaneously cutting the cost of using their models and optimizing them for long-running agentic tasks. Meta, meanwhile, is running into the next barrier: what happens when AI is asked not just to answer a user, but to actually do something in the real world.
OpenAI Makes GPT-6 Cheaper
On September 22, OpenAI expanded the GPT-6 family with GPT-6 Sol and GPT-6 Luna. The company positions them as more affordable alternatives to its flagship GPT-6 Astra for professional work, coding and AI agents.
The main bet is not only on capability, but also on price.
GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens. GPT-6 Luna goes much lower: $0.10 per million input tokens and $0.50 per million output tokens.
At the same time, OpenAI has strengthened prompt caching. Reused context can be processed at discounts of up to 90%, while developers now have tools to monitor cache performance, diagnose misses and preload relevant context.
For a regular chatbot, that is a useful optimization. For an AI agent that works for hours with the same instructions, documents and tools, it can translate into a significant reduction in operating costs.
There is another notable metric. In its internal factual accuracy testing, OpenAI says GPT-6 Sol makes roughly half as many errors as GPT-5.6 Sol. At high reasoning levels, GPT-6 Luna, according to the company, approaches the capabilities of more expensive models.
These are OpenAI’s own tests rather than an independent audit, but the direction is clear.
The company is no longer selling raw model intelligence alone.
It is selling intelligence per dollar.
Anthropic Answers With Claude Opus 5.5
Anthropic made an almost identical move.
On September 22, the company introduced Claude Opus 5.5, the first model in the new Claude 5.5 family.
API pricing is $4 per million input tokens and $20 per million output tokens, 20% lower than Opus 5. Anthropic estimates that the typical cost of completing tasks has fallen by around 40%, while generation speed has increased by more than 30%.
Cached context became even cheaper. Cache reads cost $0.20 per million tokens, 60% below the previous generation.
That matters specifically for a new class of applications.
A modern AI agent no longer answers one question and disappears. It reads repositories, launches tools and subagents, returns to earlier context and performs long chains of actions.
That makes the price of a single response increasingly irrelevant.
The real question is how much it costs to finish the entire job.
Anthropic gives a striking example. One early tester used Opus 5.5 to migrate a 680,000-line codebase. According to the company, the model completed the job in less than a day, while an engineering team could have spent weeks on the same task.
In another test, Opus 5.5 completed command-line tasks with around 40% fewer calls and roughly half the token usage compared with Opus 5.
Anthropic is also strengthening safety for autonomous work. The model now includes a mechanism for checking actions before execution, alongside stronger protection against prompt injection when using browsers, computers and external tools.
Next in line are expected to be Claude Sonnet 5.5 and Claude Haiku 5.5, extending the same strategy to more affordable models.
OpenAI and Anthropic are effectively pursuing mirror-image strategies: make powerful AI cheap enough to run continuously instead of reserving it for isolated, high-value queries.
And Then There Is Meta
Against that backdrop, Meta’s story reveals a different side of the race.
On September 22, it emerged that the company had been testing a “human concierge” function for its new AI assistant, Muse.
Muse is designed to send emails, make purchases, handle bookings and carry out other tasks. It can also call businesses on a user’s behalf — for example, to book an appointment, check whether an item is available or ask about the price of a service.
But in some cases, those calls were handed over to a live contractor.
The experiment was already operating at a meaningful scale. According to Sensor Tower data cited by Reuters, Muse had accumulated more than 2.5 million downloads and had reached the top of the US app charts.
The reason humans entered the loop was simple: people are not always willing to talk to robots.
One Meta employee said an insurance company simply ended the call once it realized it was speaking with AI.
Meta then began handing some calls to human operators. At one stage, the feature was available to roughly half of the employees participating in the test.
The result was striking: with human intervention, the success rate of calls reportedly reached 95–98%, according to Meta’s internal data.
But that improvement created another problem.
The user believes they are interacting with an AI agent and handing information to a machine. Yet at some point, a real person may enter the process.
The Main Problem Was Not Technical
That is where privacy concerns began.
If a user gives the assistant an address, booking details, account information or other personal data, some of that information could potentially be seen by the contractor handling the call.
During testing, one employee found an inappropriate racial remark in the transcript of one such conversation. Following internal criticism, a Meta executive acknowledged that the experiment had been launched without sufficient disclosure, and the feature was temporarily disabled in its original form.
Meta says such testing is necessary to identify problems before a broader release and that the technology will require stronger safeguards, privacy protections and clear user disclosure.
This is where the real limit of today’s AI agents becomes visible.
Understanding a request is no longer the hardest part.
The much harder challenge is acting reliably in the real world.
Writing an email is easier than dealing with an insurance company. Finding a hotel is easier than changing a booking. Recommending a product is easier than negotiating with a seller who asks an unexpected question.
It is precisely in that final stretch between “almost capable” and “actually completed” that humans remain extremely useful.
The Race Is No Longer About the Smartest Model
OpenAI, Anthropic and Meta are showing three parts of the same picture.
OpenAI is cutting the cost of GPT-6 and making long-context workloads cheaper. Anthropic is reducing both the cost and the number of actions required to complete complex tasks. Meta is trying to move AI agents out of the interface and into the real world — and discovering situations where the machine still has to give way to a person.
That means the new AI race is starting to look less like a competition over benchmark tables.
The real competition begins when a system must receive a task, complete it from start to finish, stay within a reasonable cost and avoid creating new problems around security and user data.
The most important metric for the next generation of AI may turn out to be very simple:
not how well the model answers, but how well it gets the job done.
Sources
OpenAI — Introducing GPT-6 Sol and Luna
OpenAI — Better prompt caching for GPT-6
Anthropic — Introducing Claude Opus 5.5
Reuters — Meta testing a “human concierge” for its new personal AI agent, Muse
