The calm that followed the launch of GPT-6 Astra — OpenAI's new flagship model, unveiled in early September and built around autonomous execution of long-horizon tasks and computer use — did not last long. Over the past few days, leaked tests and quiet beta deployments in Claude Code and Cowork suggest that Claude Fable 5.2 is not only ready for general release, but decisively outperforms Astra in high-difficulty reasoning scenarios.

1. The Context of the Showdown

In early September, OpenAI introduced GPT-6 Astra, marking a leap forward from the GPT-5.6 Sol family. Its pitch was not simply faster answers, but the complete harness engineering loop: understanding a goal, executing on the operating system, and observing and self-correcting autonomously across sessions lasting tens of minutes.

That agentic muscle, however, came with two severe trade-offs:

  • Prohibitive inference cost: A pricing scheme reported at roughly $10 USD per million input tokens and $50 USD per million output tokens in standard mode, making every task significantly more expensive than with previous models.
  • Opaque monitoring and alignment: Researchers have reported difficulty auditing reasoning chains when the model operates in long autonomous loops.

2. Where Fable 5.2 Pulls Ahead

Test cases shared by independent evaluators and leaks from the closed beta highlight three factors where Anthropic has tipped the scales:


If those signals hold up once the model ships publicly, the pressure on OpenAI will not be about raw capability, but about whether Astra's price tag can still be justified against a rival that appears to reason harder for less.