Shopping cart

Subtotal $0.00

View cartCheckout

Magazines cover a wide array subjects, including but not limited to fashion, lifestyle, health, politics, business, Entertainment, sports, science,

Productivity / Tools

GPT-6 Astra: OpenAI’s New Flagship Model Explained

GPT-6 Astra AI intelligence with computer use, agentic workflows, coding, reasoning, science and cybersecurity
Email : 6

GPT-6 Astra Model Release

OpenAI has launched GPT-6 Astra, calling it the most capable and best-aligned model it has shipped to date. Initially, the model is rolling out to a limited group of organizations, with a wider release to ChatGPT Plus, Pro, Business, and Enterprise users, plus the OpenAI API, Microsoft Azure, and AWS Bedrock, expected over the following days.

For a deeper walkthrough of what the model can do, see our complete GPT-6 Astra guide.

Specifically, Astra’s headline claim is state-of-the-art performance across computer use, web browsing, software engineering, cybersecurity, science, and general professional work, alongside major gains in staying aligned with what users actually ask for.

97.6%
FrontierMath Tier 4 (v2)
99.9%
ARC-AGI-3
100%
ExploitBench
96.0%
GPQA Diamond

GPT-6 Astra Benchmark Highlights

Astra posts some of the strongest scores OpenAI has published to date:

GPT-6 Astra vs. prior/rival models
FrontierMath Tier 4 (v2)97.6%
Similarly, this is up from 83.0–90.2% on prior models.
ARC-AGI-399.9%
In contrast, the best rival model scores just 30.2%.
ExploitBench (cybersecurity)100%
Similarly, this is up from 78.5% for GPT-5.6 Sol.
GPQA Diamond96%
In contrast, GPT-5.6 Sol scored 94.6%.
Terminal-Bench 4.0 (coding)57.9%
In contrast, rival models scored 37.3–55.8%.
Agents’ Last Exam (professional tasks)59.3%
Unlike this model, rival models scored 53.6–55.5%.

OpenAI says Astra also completes computer-use tasks roughly 47% faster than its predecessor while scoring higher, and does so using noticeably fewer output tokens than rival models in several head-to-head evaluations.

Bar chart comparing GPT-6 Astra and GPT-5.6 Sol benchmark scores
Source: OpenAI, “GPT-6 Astra: A new generation of intelligence”

GPT-6 Astra: A Stronger Computer-Use Model

In practice, Astra is positioned as OpenAI’s best model yet for operating a computer directly: filling out forms, updating CRM records, managing calendars, researching and drafting documents, and running QA checks on websites it builds. For instance, demonstrations shared alongside the launch show the model handling tasks like PCB layout in circuit design software, spreadsheet modeling, and everyday chores such as researching pediatricians or apartment listings.

Form fillingCRM updatesCalendar managementWeb researchPCB layoutSpreadsheet modelingFrontend QA

A Step Up for Professional Work

For office and creative work, OpenAI highlights Astra’s ability to match existing templates, producing slide decks, documents, and spreadsheets that follow a company’s formatting and tone rather than generic output. For example, partners such as Cognition (Devin) and Higgsfield AI report meaningfully better real-world performance, with Higgsfield citing up to 20% fewer tokens used on production workflows.

↓ 20%
Fewer tokens used on production creative workflows (Higgsfield AI)

Coding Gains

On coding-specific benchmarks, Astra leads or matches the field, including a jump on Terminal-Bench 4.0 (57.9% versus 37.3% for GPT-5.6 Sol) and strong results on FrontierCode. OpenAI is also rolling out a note-taking mechanism in Codex that lets Astra retain context across long sessions without repeatedly compressing earlier work into summaries.

Terminal-Bench 4.057.9%
GPT-6 Astra
Terminal-Bench 4.037.3%
GPT-5.6 Sol

Advancing Science

In fact, OpenAI says Astra contributed to two new mathematical results on the distribution of prime numbers, and it posts strong scores across science and health benchmarks, including GPQA Diamond (96.0%) and HealthBench Professional (63.4%).

96.0%
GPQA Diamond
63.4%
HealthBench Professional

Cybersecurity: More Capable, More Guarded

Meanwhile, Astra’s jump in cyber capability is significant enough that OpenAI classifies it as “Critical” under its own risk framework. In internal testing without production safeguards, it reached a 100% success rate on ExploitBench and discovered two previously unknown zero-day vulnerabilities, which OpenAI is disclosing to the affected maintainers. In production, the model is restricted from generating proof-of-concept exploits, with looser restrictions planned in an upcoming Daybreak program.

⚠ Critical risk classification

Consequently, under OpenAI’s Preparedness Framework, Astra meets the Critical threshold for cybersecurity — the company is pairing the release with tighter production safeguards even as raw capability jumps sharply.

ExploitBench — GPT-6 Astra100%
ExploitBench — GPT-5.6 Sol78.5%

Alignment and Safety

Overall, OpenAI frames Astra as its most aligned model to date, saying it is far less likely to go beyond the scope of a task or misrepresent its own capabilities compared with GPT-5.6 Sol. The company is pairing the release with expanded misalignment monitoring in production.

GPT-6 Astra Pricing and Availability

Astra will be available through the OpenAI API under the model name gpt-6-astra, as well as via Microsoft Azure and AWS Bedrock. Additionally, standard API pricing is $10 per million input tokens and $50 per million output tokens, with a faster processing tier available at double that price. However, enterprise admins will need to switch it on for their workspace, since it is off by default at launch.

Model name (API)
gpt-6-astra
Input tokens
$10 / million
Output tokens
$50 / million
Fast mode
2× price, 2× speed
Also available via
Microsoft Azure, AWS Bedrock
Enterprise default
Off — admins must enable

Source: OpenAI, “GPT-6 Astra: A new generation of intelligence”

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts