• Home
  • Fresh Content
  • Courses
  • Resources
  • Podcast
  • Talks
  • Publications
  • Sponsorship
  • Testimonials
  • Contact
  • Menu

Jon Krohn

  • Home
  • Fresh Content
  • Courses
  • Resources
  • Podcast
  • Talks
  • Publications
  • Sponsorship
  • Testimonials
  • Contact
Jon Krohn

OpenAI’s GPT-6 Astra

Added on September 15, 2026 by Jon Krohn.

At GPT-6 Astra's launch, OpenAI president Greg Brockman suggested future observers may point to it as the arrival of AGI. In fact, Artificial Analysis's latest benchmarking suggests it merely catches OpenAI up to Anthropic. Here are the key details:

GPT-6 ASTRA BASICS
• Successor to GPT-5.6 Sol, trained on 100,000+ GPUs at the Stargate facility in Texas.
• It's the first OpenAI model where other A.I. models helped supervise training.
• API pricing is $10/M input tokens and $50/M output (~2.5x Sol's promotional rate; in line with Claude Fable 5.1), with five reasoning-effort settings from low to max.
• Rolled out to ChatGPT Plus, Pro, Business and Enterprise (Enterprise admins must switch it on).

GPT-6 ASTRA CAPABILITIES
• Computer use: 72.6% on OSWorld 2.0 (vs 65.7% for Sol) while finishing tasks ~47% faster.
• Coding: ~58% on Terminal-Bench 4.0 (vs 37% for Sol, ~56% for Fable 5.1), though level with the top Claude models on FrontierCode.
• Math: 99.9% on ARC-AGI-3 (Sol scored under 8%) and 97.6% on FrontierMath Tier 4. Astra also helped tighten a bound on prime gaps from 240 to 186.
• Science: ~65% on Terminal-Bench Science (vs ~22% for Sol, ~53% for Fable 5.1) and ~65% fewer output tokens than Claude Opus 5 on Agents' Last Exam.
• *All of the above stats, however, come from OpenAI themselves.* The third-party view (see chart) shows Astra at 53, tied with Claude Fable 5.1 (and up from Sol's 47) on Artificial Analysis's (widely respected) composite "Intelligence Index".

SAFETY
• First OpenAI model rated "Critical" for cybersecurity under their Preparedness Framework: without safeguards it finds unknown vulnerabilities and builds working exploits (100% on ExploitBench vs 78.5% for Sol; two zero-days found in a fresh Chrome evaluation).
• The release was slowed and gated. The public version supports defensive work but refuses exploit writing; vetted orgs in OpenAI's "Daybreak" program get fewer restrictions.
• New misalignment monitors can pause or halt actions that look unauthorized.
• Encouraging: on a scope-creep evaluation, Sol exceeded its authorized target 48% of the time; Astra did so 0% of the time.
• Concerning: Astra's written reasoning is harder to monitor than Sol's.

IS IT AGI?
• AGI isn't a binary event; there are degrees of breadth and depth (listen to Episode #748 of my podcast for more details).
• BOTTOM LINE: Astra looks like a noteworthy jump in general-intelligence capability for OpenAI relative to their previous flagship model (5.6 Sol), but it does not appear to be moving the frontier toward AGI in any major way... it is merely catching up to Anthropic at the frontier vicinity.

The SuperDataScience podcast is available on all major podcasting platforms, YouTube, and at SuperDataScience.com.

In Data Science, Five-Minute Friday, Podcast, SuperDataScience, YouTube Tags #superdatascience, #GPT6, #OpenAI, #anthropic, #ai, #agi
Older: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano →
Back to Top