# Four Frontier Labs Shipped New Models in the Same Week. Here's What Actually Changed.

OpenAI, Anthropic, Google, and Meta all released new models in the first days of September 2026 - and xAI followed three weeks later. Most of what shifted is pricing and API compatibility, not the headline capability claims.

By Priya Nair (AI Industry Reporter) — published 2026-09-15, updated 2026-09-22
Source: https://aiscoutdaily.com/news/four-frontier-labs-shipped-models-same-week

## TL;DR
OpenAI's GPT-6 Astra (Sept 3-4), Anthropic's Claude Fable 5.1 and Mythos 5.1 (Sept 1), Google's Gemini 3.8 Flash (Sept 2), and Meta's Muse Spark 1.3 (Sept 2) all shipped within days of each other, with xAI's Grok 4.7 following September 21. GPT-6 Astra's benchmark claims (99.9% on ARC-AGI-3) are the most striking on paper, but the benchmark's own creators have pushed back on treating a high score as evidence of general intelligence. For most people picking a chatbot, the practical changes this week are pricing and API compatibility, not a fundamental capability jump.

If it felt like every AI lab announced something at once this month, that's because they did. OpenAI, Anthropic, Google, and Meta all shipped new models within the first four days of September 2026, and xAI followed three weeks later with Grok 4.7 on September 21. Fourteen new models shipped from ten providers over the month - the kind of release cadence that makes "which AI should I use" a genuinely different question every few weeks.

## What actually shipped

- OpenAI's GPT-6 Astra launched as a limited preview on September 3, with broad availability the next day across ChatGPT Plus, Pro, Business, and Enterprise, plus the API, Azure, and AWS Bedrock.
- Anthropic shipped Claude Fable 5.1 and Mythos 5.1 on September 1, at an unchanged list price but with three breaking API changes worth checking if you're building on top of Claude.
- Google released Gemini 3.8 Flash on September 2, at the same introductory price as the 3.7 Flash it replaces.
- Meta released Muse Spark 1.3 on September 2, adding a contributor tier.
- xAI's Grok 4.7 followed on September 21, the most recent of the wave.

## The benchmark claim worth a second look

GPT-6 Astra's headline numbers are the most striking of the batch: OpenAI reports 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench. That ARC-AGI-3 number in particular has been read by some as evidence the industry has crossed into "the AGI era." ARC Prize, the organization that built the benchmark, has been explicit that this reading goes further than the evidence supports - its own stated position is that saturating the benchmark "would not represent proof of achieving AGI."

That distinction matters for anyone actually choosing a tool: a model topping one benchmark, even a hard one, doesn't mean it's the best choice for your specific use case. Terminal-Bench 4.0, a benchmark more directly relevant to coding work, currently has Claude Fable 5.1 and GPT-6 Astra essentially tied for first (57.9% and 58.2% respectively) - a much closer race than the AGI framing suggests.

## What this means if you're picking a chatbot today

For most day-to-day use, the practical differences this week are pricing and API compatibility, not a fundamental capability jump. If you're on Claude and rely on the API directly, check Anthropic's breaking-change notes before you upgrade. If you're choosing between GPT-6 Astra and Claude Fable 5.1 for coding work specifically, the two are close enough on real benchmarks that workflow fit and pricing should probably decide it, not the AGI headlines.

## Sources
- [Four major AI labs launch new models in the first week of September 2026](https://www.dutchstartup.ai/en/news/four-major-ai-labs-launch-new-models-in-the-first-week-of-september-2026)
- [AI Model Releases: September 2026 Tracker and Dated Ledger](https://www.digitalapplied.com/blog/ai-model-releases-september-2026-tracker)
