MeshioMeshio
News

OpenAI’s GPT-6 Astra targets long-context agent workloads at $10/$50

GPT-6 Astra brings a 1.05-million-token context window, new reasoning levels and a staged rollout—but its benchmark claims remain unverified.

Meshio Newsroom
Meshio NewsroomSep 4, 2026

OpenAI has announced GPT-6 Astra, a flagship model aimed at demanding coding, research, computer-use and document workflows. LLM Stats reports that the model entered staged general availability on September 3, 2026, with Trusted Access enterprise customers first and broader access expected in the following days.

Astra’s headline capability is a 1,050,000-token context window, paired with up to 128,000 output tokens. It accepts text and images and produces text, while supporting tools including web search, file search, code execution, hosted shells, computer use, MCP and structured outputs. Fine-tuning is not supported.

Pricing and controls

Standard API pricing is $10 per million input tokens and $50 per million output tokens. Cached input costs $1, while cache writes cost $12.50. Requests exceeding 272,000 input tokens incur a 2× input/cache multiplier and a 1.5× output multiplier. Batch and Flex pricing is half the standard rate; Fast mode costs twice as much.

The API exposes five reasoning settings—low, medium, high, xhigh and max—but defaults to low. Builders seeking the model’s highest reasoning setting will need to specify it explicitly, which could affect latency and cost.

Benchmark claims need context

OpenAI’s self-reported launch table lists scores including 57.7 on Terminal-Bench 4.0, 74.1 on DeepSWE v1.1, 64.6 on Terminal-Bench-Science, 91.5 on BrowseComp and 72.6 on OSWorld 2.0. LLM Stats has not independently verified those results, and notes that OSWorld is an offline partial evaluation while several ARC and GPQA results are near ceiling.

For AI developers, Astra’s combination of unusually long context, extensive tools and configurable reasoning could simplify complex agent pipelines. But staged availability, the low default effort setting, usage multipliers and unverified benchmarks mean teams should test real workloads before committing.

Source: LLM Stats

Comments

Log in to join the discussion