OpenAI claims we've entered the AGI era — has GPT-6 Astra really demonstrated general intelligence?
OpenAI representatives claim that the recent launch of one of its latest models, GPT-6 Astra, marks the beginning of the artificial general intelligence (AGI) era — a hypothetical scenario in which artificial intelligence (AI) can learn and reason like humans.
Following very recent claims from technology executives that we've now reached this milestone, how likely is this claim to stand up to scrutiny? Amid revelations that safety concerns have forced the company to abandon the planned rollout of GPT-6.1 Astra , independent researchers and benchmark creators warn that GPT-6 Astra's record-breaking scores rely on heavy prompt optimization rather than true general intelligence.
This comes as experts warn about the risks of future AI systems and AI executives jointly call for a slowdown in AI research .
In announcing the model's release, OpenAI President Greg Brockman suggested that Astra signaled a fundamental turning point in AI evolution.
"If we fast-forward a couple years, and we look back and say when was it really that AGI was created, I think it's going to be about this time, and I think it might be about this model," Brockman said during a Sept.
3 news conference marking the launch.
How did the new model perform in benchmarking? The benchmarks scores are impressive, but even the creator of one of the most important suggests acing it doesn't automatically mean we've reached AGI. (Image credit: Cheng Xin via Getty Images) As part of the testing-and-verification process for the new model, OpenAI released benchmark results across a range of tests that measure the capabilities of AI models at various tasks.
OpenAI representatives claimed in a statement that the company's new model delivers "state-of-the-art" performance across a variety of fields, including software engineering, the autonomous use of computer programs, mathematics problems, and even scientific research.
Demonstrations, for example, highlighted Astra's ability to generate 3D CAD code; lay out printed circuit boards in CAD (computer-aided design) software; convert digital 3D models into interactive, video-game-style environments; and deploy hosted web applications directly from prompts.
"In our early testing, Astra stood out by approaching legal work the way a discerning lawyer does: it distinguishes documents from established records, surfaces unsupported assumptions, and converts gaps into concrete drafting positions," said Niko Grupen , head of applied research at legal services AI developer Harvey as part of OpenAI’s announcement.
On ARC-AGI-3 — a benchmark designed to measure how efficiently AI systems acquire new skills in novel, abstract environments — Astra achieved a headline score of 99.9%.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.livescience.com — the content belongs to Live Science.