By Evan Vega
SAN FRANCISCO — The release of OpenAI’s GPT-6 Astra has ignited a national debate over the actual state of artificial general intelligence (AGI), as critics argue that viral marketing and influencer commentary are outpacing the company’s technical disclosures.
Launched Sept. 3, GPT-6 Astra demonstrates significant technical leaps in reliability. According to OpenAI, scope violations dropped from 48% to zero, and hallucinations fell from 12.2% to 4.2%. Falsified labels also saw a sharp decline, moving from 36 per 100 to 17 per 10,000.
Despite these gains, the company has avoided explicitly claiming the achievement of AGI. While OpenAI co-founder Greg Brockman stated it is “not unreasonable to feel that we are now in the AGI era,” the company’s official launch materials stop short of a definitive declaration.
The friction lies in the interpretation of benchmark data, specifically the ARC-AGI-3. Viral content, including videos from the “AI News & Strategy Daily” channel which have garnered approximately 400,000 views, has highlighted a 99.9% success rate. However, that figure stems from a “Provider Adapter” run—a specialized harness designed by OpenAI that combines the model with specific tools.
On the standard harness used to score all other competing models, GPT-6 Astra scored 62.7% at maximum effort. Data indicates that the adapter harness contributes significantly to the score; when reasoning is disabled, Astra scores 96.7% on the adapter but only 35.2% on the standard harness.
Industry analysts note that these distinctions are documented in OpenAI’s technical footnotes and a July 29 report titled “How enabling two settings tripled our scores on the ARC-AGI-3 benchmark.” The report acknowledged that benchmarks often measure “less visible choices about API settings, harness design, and prompting.”
Further complexities appear in the model’s comparison to competitors. Footnotes in the launch materials reveal that some high-performance scores were achieved using “Mythos,” a version of the model with fewer safeguards. This version is restricted to vetted organizations, meaning the versions of the software available to the general public may not reflect the peak performance cited in marketing tables.
The discrepancy highlights a growing gap between the technical reality of large language models and the public perception of autonomous AI capable of handling complex, week-long assignments without human instruction.
Related: Frontier Watch