
GPT-6 Astra: What It Costs and What It Can Automate
GPT-6 Astra costs $10 per million input tokens, holds a 1.05M-token context, and finished 72.6% of OSWorld desktop tasks in roughly half the time.
OpenAI's Newest Model Is Built to Operate Software, Not Just Discuss It
OpenAI began rolling out GPT-6 Astra on September 3, 2026, and the most consequential detail is not a chat improvement. It is that the model is designed to drive a computer the way a person does — moving through browsers, spreadsheets, desktop applications and web forms to finish a task end to end. The company positions Astra as its strongest model for software engineering to date, and the published numbers that matter most sit in the computer-use column rather than the conversation column.
- Pricing: $10 per million input tokens, $1 per million cached input tokens, and $50 per million output tokens, with higher rates on prompts above 272,000 tokens
- Context window: 1,050,000 tokens total — 922,000 maximum input and 128,000 maximum output — with an April 30, 2026 knowledge cutoff
- Computer use: 72.6% on OSWorld 2.0 at roughly 40 minutes per task, against GPT-5.6 Sol's 65.7% at roughly 75 minutes
- Access: enterprise Daybreak customers first, then ChatGPT Plus, Pro, Business and Enterprise, plus the OpenAI API and Amazon Web Services
What Does GPT-6 Astra Cost to Run?
OpenAI lists the model as gpt-6-astra at $10 per million input tokens and $50 per million output tokens, with cached input dropping to $1 per million. Prompts that run past 272,000 tokens move to a higher tier. Inside ChatGPT, Astra draws on existing subscription allowances rather than carrying a separate line item, and Pro, Business and Enterprise subscribers get access to a higher-effort variant, GPT-6 Astra Pro.
The cached-input rate is the number worth planning around. A tenfold discount on repeated context is not a rounding detail when the model is meant to hold a codebase, a document set or an application's state across a long-running job. Combined with the selectable reasoning effort levels — low through max — the pricing sheet reads as one built for sustained agentic work rather than short exchanges, which is a meaningful shift from how these models were priced two generations ago. Readers who followed OpenAI's 80% price cut on Luna will recognise the pattern: capability at the top of the range, aggressive discounting on the repeated work underneath it.
What Can Astra Actually Do on a Desktop?
OpenAI describes Astra filling out forms, updating CRM records, running research and drafting summaries, analysing data, building websites, performing front-end QA passes and troubleshooting problems visible on screen. It also produces finished documents, spreadsheets and presentations rather than the raw text a human then has to assemble.
The OSWorld 2.0 result is the clearest evidence behind those claims. Astra completed 72.6% of tasks in that desktop benchmark at roughly 40 minutes per task, where its predecessor GPT-5.6 Sol managed 65.7% at roughly 75 minutes. Accuracy up, wall-clock time close to halved. For anyone evaluating whether an agent can be trusted with a multistep workflow, that second figure is arguably the more practical one — a capable agent that takes an hour and a quarter per task is a very different operational proposition from one that takes forty minutes.
Codex Keeps Notes Instead of Summaries
The Codex side of the release carries a quieter but genuinely useful change: a context-preservation feature that maintains searchable notes across sessions rather than compressing prior work into a summary. Summarisation is lossy by construction, and the loss tends to fall on exactly the specific details — a file path, a flag, a decision and its reason — that a coding agent needs later. Retrievable notes sidestep that failure mode.
How Much Weight Should the Benchmark Numbers Carry?
OpenAI reports 97.6% on FrontierMath Tier 4 v2, 96% on GPQA Diamond, 74.1% on DeepSWE v1.1 and 100% on ExploitBench. Reported figures in the mid-to-high nineties on a benchmark generally mean the benchmark is approaching the end of its useful life, not that the remaining questions were easy. That is a real result and also a measurement problem: once a test saturates, it stops discriminating between models, and the field needs harder instruments before the next comparison means much.
These are the vendor's own numbers on the vendor's own harness, and no independent lab has published a comparable run yet. That is the standard caveat for every frontier release, and it applies here in full. The ExploitBench score in particular sits behind the staged access programme we covered in OpenAI's cyber safeguards rollout — those capabilities are not part of general availability. Astra has been in view for a few weeks already; it is the same model that proved ten open math problems in Lean 4 in August.
OpenAI leadership has framed the launch in AGI terms, with president Greg Brockman saying "I think it's not unreasonable to feel that we are now in the AGI era." That is a claim about interpretation, not a measured result, and it is worth reading as such. Chief scientist Jakub Pachocki paired it with a more concrete commitment — "progress in intelligence does not guarantee progress in alignment" — which is the sentence the rest of our AI coverage will be checking against over the next few months.
The grounded read is simpler. A model that finishes more desktop work correctly in half the time, holds a million tokens of context, and discounts repeated context tenfold changes what is economically sensible to automate. That is a substantial step, and it does not need a bigger word attached to it.
Sources: OpenAI API documentation — September 3, 2026; 9to5Mac — September 3, 2026; Fox Business — September 3, 2026; VentureBeat — September 3, 2026.
More AI Stories

WeatherNext 3 Brings 5km Hourly Forecasts to Search
Google DeepMind's WeatherNext 3 forecasts at 5km resolution every hour and improves rain accuracy up to 60% over WeatherNext 2, live in Search now.

Gemini Agentic Video Understanding Cuts Tokens 88%
Gemini agentic video understanding cuts token use up to 88% and cost up to 66% while raising accuracy 7%, with no extra fee on the Gemini API.

Gemini 3.8 Flash Cyber Fixes 2.6x More Chrome Bugs
Google's Gemini 3.8 Flash Cyber wrote 2.6x more correct Chrome patches than larger commercial models, and 3.8 Flash starts at $0.75 per million tokens.
