6 articles tagged “benchmarks”, most recent first.
GPT-6 Astra brings computer use, coding, browsing and document reasoning to existing business workflows, with API pricing from $10 per million input tokens.
GPT-6 Astra targets autonomous computer work, software engineering, science, and safer delegation with major benchmark gains.
GPT-6 Astra brings a 1.05-million-token context window, new reasoning levels and a staged rollout—but its benchmark claims remain unverified.
Qdrant and Vultr have released a free 10-billion-vector dataset and open-source tooling for testing retrieval systems at internet scale.
A controlled experiment suggests familiar human knowledge may explain much of the behavior seen in large multi-agent worlds.

Google DeepMind is testing a cryptographic way to keep both model weights and evaluation prompts hidden during independent AI assessments.