
OpenAI's GPT-6 Astra AGI Claim Undercut by ARC Prize's Own Benchmark Verification
The AMW Read
Independent benchmark re-verification disputes a top-lab AGI claim and directly tests the scaling-bull versus skeptic frames over whether headline scores reflect true model capability.
OpenAI's GPT-6 Astra AGI Claim Undercut by ARC Prize's Own Benchmark Verification
OpenAI unveiled GPT-6 Astra on September 3 as "the world's most intelligent, safest model," with co-founder Greg Brockman calling the result on social media "this feels different — I think this is AGI." OpenAI's own benchmark data cited a 99.9% score on ARC-AGI-3, the novel-puzzle reasoning test built by the ARC Prize Foundation specifically to resist memorization. But ARC Prize's independent verification found that figure only holds inside OpenAI's custom "harness" — tooling that lets the model retain its full reasoning trace across game resets rather than starting fresh each round. Under the standard harness used to compare all models on equal footing, Astra scored 62.7%, a 37-point gap. ARC Prize co-founder Mike Knoop and outside researchers argue the 99.9% number measures the surrounding tool stack, not the model itself, and note that Anthropic's Claude Opus 5 scored only 30.2% under the same optimized harness — evidence that scaffolding, not raw model capability, drove Astra's headline result.
The dispute lands on a fault line the market has argued since benchmark scores became marketing collateral: whether a lab's own harness-optimized number is a valid AGI proxy or a way to launder tool-assisted performance into a capability claim. ARC-AGI-3 was built to block memorization at the training-data level, but ARC Prize's finding shows the same gap reappears one layer up, at the harness level, where model-card headlines rarely disclose which configuration produced the score.
For builders, the 62.7% standard-harness result — not 99.9% — is the realistic baseline for production systems that won't ship with OpenAI's proprietary memory-retention scaffolding attached. For investors, ARC Prize's plan for a less-predetermined ARC-AGI-4 in 2027 signals that benchmark methodology itself is still a moving target, and any lab's AGI claim should be treated as provisional until harness-controlled, third-party verification exists.



