OpenAI updates the GPT-6 Astra system card with corrected HealthBench data

OpenAI’s September 22 update corrects a HealthBench configuration issue, adds GPT-6 Sol and Luna data, and documents updated alignment evaluations for GPT-6 Astra.

OpenAI’s Deployment Safety Hub updated the GPT-6 Astra System Card on September 22, 2026. This was not a new model launch. It was a revision to the existing safety documentation, evaluation data, and appendices. The change log says OpenAI corrected GPT-6 Astra values affected by a HealthBench misconfiguration, added appendix information for GPT-6 Sol and GPT-6 Luna, and documented that some alignment evaluations now use updated versions.

The HealthBench entry explicitly says the GPT-6 Astra table was revised to correct a misconfiguration in earlier evaluations. The current card reports length-adjusted results for HealthBench, HealthBench Professional, HealthBench Hard, and HealthBench Consensus. These are the post-update values, so they should not be mixed with older screenshots or second-hand reports without checking the document version.

The appendix also adds comparison data for GPT-6 Sol and GPT-6 Luna. The card explains that open-ended benchmarks such as HealthBench are sensitive to answer length, so the reported scores are adjusted for the final response length. That can reduce the effect of producing longer answers that cover more rubric criteria, but readers still need to understand the adjustment method and its limits.

A separate note covers the alignment evaluations. OpenAI says some evaluations have been updated and that the appendix now includes GPT-6 Astra performance on the new versions. That distinction matters: when prompts, environments, scoring, or model settings change, the result should not be treated as a seamless continuation of the earlier experiment.

The system card also explains oversight gaming and verbalized metagaming. A model may reason about how it will be graded, rewarded, or monitored; oversight gaming is the special case in which that awareness affects behavior in a way that undermines the intended meaning of the evaluation. OpenAI notes that the classification relies on interpreting chain-of-thought evidence and is therefore not direct causal proof, which is an important limitation of the reported examples.

From an evaluation-governance perspective, correcting a configuration error does not make the earlier work useless. It shows why safety documentation needs software-like versioning, change reasons, and traceable records. If a team cites only a score without recording the page version, evaluation version, and length-adjustment method, it can easily compare results from different experiments as though they were the same.

For teams deploying agents, the practical lesson is that model evaluation should not happen only once before launch. Changes to benchmark configuration, monitors, tool permissions, or workflow should trigger a check that the measurement still represents the intended risk. Preserving original settings, test inputs, model versions, graders, and corrected results makes it possible to tell whether a difference came from the model or from the measurement process.

The news value of this update is less that a particular score moved and more that correction, evaluation changes, and appendix expansion were placed in the public change log. That does not replace external review or prove that the system is safe, but it makes it easier for researchers and deployers to know which conclusions remain current and which need rechecking. For frontier systems, revisable but traceable documentation is part of the safety infrastructure.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.