Unconfirmed: the first developer verdicts on Anthropic's Claude Opus 5.5 are landing, and they disagree. Automatica has not reproduced any of the results below, and none of the accounts published a methodology.
Bindu Reddy, chief executive of Abacus.AI, posted an unfavourable comparison: the model "simply requires a lot of prompt iteration compared to Fable", with a "Net result - 2x my Fable bill not to mention my time", and said she was "Going back to Fable 5.1". That is one workload, self-reported, with no task list or token counts attached.
Matt Shumer has been posting a different kind of observation, about a long-running project he says the model keeps extending on its own. In one post he wrote that it built "a weather system that matches whatever weather is in NYC at the time", so during a nor'easter "there's rain, and the NPCs are carrying umbrellas". In another he said the model "decided to add helicopters" and "tells me it's going to keep making them better throughout the day."
The two accounts are not in conflict so much as measuring different things: one is counting money against a competitor on a specific workload, the other is describing behaviour that is hard to price. Both are worth reading as impressions rather than evidence, which is how their authors present them.
The model itself is real and generally available: Opus 5.5 arrived on Amazon Bedrock and Perplexity this week, which our news desk reported from the companies' own announcements.
What would settle the cost question is dull and public: the same task set run against both models with token counts and prices published. Several developers have said they will do exactly that. Until then, "2x the bill" is one person's invoice, not a benchmark.