$ 44.57 € 51.77 zł 12.01
+15° Kyiv +14° Warsaw +24° Washington

OpenAI changed GPT-6 Astra model evaluation metrics — Fortune

Lev Shevtsov 05 September 2026 03:46
OpenAI changed GPT-6 Astra model evaluation metrics — Fortune

On September 3, U.S. Eastern Time, OpenAI changed the model test results listed in its GPT-6 Astra announcement after its publication. Some of the adjustments temporarily improved its position compared with competitors’ developments, Fortune reports, citing archived versions of the page and company comments.

Publication delay

OpenAI planned to publish the post on September 3 at 2:00 p.m. U.S. Eastern Time, but the page became available about two hours later. The company initially attributed the issue to a content management system error, and later to internet disruptions. It also told Fortune that it withdrew the first version of the publication for undisclosed reasons which, according to the company, were unrelated to the test results.

According to archived copies, Astra’s stated hallucination rate initially stood at 4.2%, but fell to 2% in a version of the page saved at 5:20 p.m. The figure for the previous GPT-5.6 Sol model was changed from 12.2% to 9.4% during the same period. At the time Fortune prepared its article, both values had returned to 4.2% and 12.2%, respectively.

More current news is available on the UA.News Telegram channel Telegram.

Changes in model evaluations

In the FrontierMath Tier 4 (v2) test, Astra’s result remained at 97.6%, but the scores for GPT-5.6 Sol and Anthropic’s Fable 5.1 were changed. Fable 5.1’s score initially stood at 87.8%, then fell to 78%, and was later adjusted to 83%. Sol’s figure changed from 83% to 80.5%, and subsequently returned to 83%.

OpenAI also changed Sol’s result in an internal version of the ExploitBench cybersecurity test from 5.5% to 11.5%. The company said it was considering restoring the previous number because the higher result reflects a level of reasoning that is not commercially available for Sol.

An OpenAI representative told Fortune that evaluation results can differ by several percentage points depending on the model version, toolset, and specific test run. According to the company, the changes in the announcement were intended to reflect its best estimate of available performance. Researchers at Stanford’s Intelligent Systems Laboratory and Center for Research on Foundation Models noted that rerunning tests under different conditions can be beneficial for marketing. At the same time, Snorkel AI engineer Vincent Sun Chen said that updating metrics ahead of a model launch is not unusual because of refinements to the model version, configuration, computing resources, and evaluation methodology.

Read us on Telegram and Sends

Download our app