GPT-6 Astra Performance and Benchmark Overview

Chase AIgo watch the original →

OpenAI's GPT-6 Astra introduces significant gains in coding, cyber security, and context management, offering competitive performance against Anthropic's Fable 5.1 at a lower price point.

Performance and Benchmarks

GPT-6 Astra demonstrates substantial improvements over the 5-series models across professional coding and security benchmarks. In Terminal Bench 4.0, the model improved from a 37.3 score in GPT-5.6 Soul to 57.7. The Deep Suite benchmark, which evaluates long-running agentic tasks, shows a 74.1% accuracy rate. Security capabilities also saw marked gains, with Exploit Bench scores rising from 55.9% to 88.5%. The model maintains high accuracy in long-context retrieval, scoring 100% on the 8-needle test for 256K to 512K tokens and 96.3% for contexts up to 1 million tokens.

Operational Efficiency and Features

While higher effort levels (Max) can lead to diminishing returns in accuracy versus cost, GPT-6 Astra remains highly token-efficient compared to Fable 5.1. For instance, at a high effort level, GPT-6 Astra achieves 57.9% accuracy for $7.21, whereas Fable 5.1 reaches 49.4% accuracy at $10.50.

Key functional updates include:

  • Context Management: The model implements a new auto-compaction feature that retains searchable notes across sessions rather than relying on a single summary document. This feature is currently experimental and requires manual activation in the config.
  • Visual Judgment: OpenAI claims improved aesthetic judgment for rendering, web design, and 3D modeling, though performance remains subject to prompt quality.
  • Hallucination Reduction: The model reports a decrease in hallucination rates from 9.4% in the 5.6 series to 2.0%.
  • Safety: Similar to Fable, Astra includes built-in guardrails that refuse to generate content identified as malicious exploits.

Pricing and Availability

GPT-6 Astra is priced at $10 per million input tokens and $50 per million output tokens. It is currently rolling out to a limited set of organizations with general API access following shortly.

  • #news
  • #ai
  • #dev-tooling

summary by google/gemini-3.1-flash-lite. probably wrong about something. check the source.