Beta The Briev beta is out. Free on iPhone via TestFlight — install it in under a minute.

Join the beta ↗
Briev
Live
Technology

OpenAI revises Astra benchmark scores amid rollout glitches and competitor comparisons

OpenAI altered several evaluation metrics for its GPT-6 Astra model after a delayed blog post launch, with some figures improving for Astra and others worsening for rival Anthropic models.

OpenAI’s announcement of the GPT-6 Astra model on Sept. 3 experienced a rollout problem, with the blog post delayed and initially inaccessible due to technical errors. After the page became viewable, the firm posted revised evaluation numbers that made Astra appear stronger on several benchmarks, including a halved hallucination rate and higher math scores, while Anthropic’s competing model saw its scores drop. The company later adjusted the same metrics again, restoring earlier values for hallucination rates and tweaking other figures such as cybersecurity and coding performance.

OpenAI said the changes reflect “best-estimate” conditions and acknowledged that evaluation noise can vary by checkpoint and setup. Critics from Stanford and Snorkel AI warned that such rapid re-running of tests could amount to “benchmaxxing,” complicating the interpretation of leaderboard results that influence market perception and investor confidence.

Why it matters

Benchmark tweaks affect how AI models are compared, influencing customer choices and investor confidence in a fast-moving industry.

In this story

benchmark scoresGPT-6 Astrahallucination rateAnthropic Fable 5.1evaluation metricsAI model comparisonbenchmaxxingOpenAI blog rollout
Get the beta ↗