GPT-5.6 Goes Live: Three Models, a New Price Floor, and a Fight Over Who Gets to Ship
OpenAI moved GPT-5.6 from a government-restricted preview to full general availability in thirteen days. What shipped, what it costs, the benchmark asterisks, and how it lands against Claude and Gemini.

OpenAI moved its GPT-5.6 family from a locked-down preview to full general availability in thirteen days. On July 9, 2026 the models went live across ChatGPT, Codex, and the API, closing a launch that had started under an unusual government-brokered restriction. This is the broad picture: what shipped, what it costs, what changed in the fine print, and how it lands against Claude and Gemini.
What shipped (the facts)
GPT-5.6 is three models under one release, tiered by capability and cost:
| Model | Positioning | API price (input / output, per 1M tokens) |
|---|---|---|
| Sol | Frontier — hardest coding, research, cyber, science | $5.00 / $30.00 |
| Terra | Balanced — everyday knowledge work | $2.50 / $15.00 |
| Luna | Fast, cost-sensitive, high-volume | $1.00 / $6.00 |
Shared specs across all three:
- 1,000,000-token context window
- 128,000 max output tokens
- Knowledge cutoff: February 16, 2026
- Prompt caching supported on all tiers (discount rate not disclosed at launch)
The headline new feature is Ultra — a highest-capability setting on Sol that coordinates multiple agents across parallel workstreams to finish long tasks faster. The API also gained programmatic tool calling (JavaScript orchestration of tool calls), explicit prompt-cache breakpoints, and image detail control (detail: original).
Availability rolled out globally over roughly 24 hours from the July 9 announcement, spanning ChatGPT, Codex, and the OpenAI API.
The benchmarks — and the asterisks
OpenAI’s pitch is frontier intelligence at lower token cost. Its reported figures:
- Agents’ Last Exam (long-running professional workflows across 55 fields): Sol 53.6, which OpenAI frames as ~13 points ahead of Anthropic’s frontier model, with Terra/Luna also beating it “at a fraction of the cost.”
- Terminal-Bench 2.1 (agentic CLI): Sol Ultra 91.9%, Sol 88.8%, versus ~88.0% for the nearest competitors.
- SWE-Bench Pro: Sol 64.6%.
The asterisks matter, and we state them plainly:
- On SWE-Bench Pro, independent write-ups report Anthropic’s frontier model scoring well above Sol (~80%) — a benchmark OpenAI publicly criticized rather than topped.
- Independent reviewer Simon Willison notes ~30% of SWE-Bench Pro tasks are considered broken, and that Sol did not clearly outperform the leading Claude model on his own complex-coding trials. [speculation] Vendor-reported eval leads should be read as marketing until reproduced.
- Competitor names and scores are inconsistent across secondary sources (some cite “Fable 5,” others “Mythos 5”). Treat cross-vendor comparison numbers as directional, not settled.
The rocky road to GA (the story CEOs miss)
GPT-5.6 did not launch cleanly. On June 26, 2026, OpenAI limited the release to “a small group of trusted partners” at the request of the Trump administration, which cited safety concerns under an executive order requiring advance government review of frontier models (up to 30 days pre-launch).
OpenAI made its displeasure public: “We don’t believe this kind of government access process should become the long-term default,” arguing restrictions keep the best tools from “users, developers, enterprises, cyber defenders, and global partners.” It framed the preview as a short-term step while it worked with the administration on a repeatable cybersecurity review process. Thirteen days later, the gate opened to everyone.
Context: this is part of a broader squeeze on frontier labs — when Anthropic released its frontier model, the same administration ordered access removed for foreign nationals, and the company pulled the model offline entirely.
The new usage conditions (what the fine print actually changed)
The ITmedia headline flagged “new usage conditions,” and there are two distinct layers:
1. Standing terms. Usage remains governed by OpenAI’s Terms of Use and Usage Policies: no illegal/harmful/abusive use, and — notably — no circumventing rate limits or bypassing safety mitigations. GPT-5.6 ships with in-model protections plus real-time monitoring; some higher-risk biology and cybersecurity requests may be refused or gated behind extra checks. Under OpenAI’s Preparedness Framework, all three models are treated as “High” capability in Cybersecurity and Biological/Chemical risk, but none reaches the “Critical” threshold.
2. A demand-driven relaxation. On July 12, 2026, after what OpenAI’s product lead called an “intense” 48 hours of Codex and ChatGPT Work usage, OpenAI temporarily removed the five-hour usage restriction for Plus, Pro, and Business plans and issued a one-time usage reset for all users. Alongside it, the model was made more efficient, so it “consumes less of your available usage.” This is explicitly temporary; no permanent terms-of-service rewrite accompanied it. (Free/Enterprise treatment was not specified.)
Where it sits in the market
GPT-5.6 resets the frontier price floor: Sol’s $5 / $30 output rate undercuts Anthropic’s frontier tier while OpenAI claims a token-efficiency edge on top. The competitive shape as of mid-July 2026:
- Anthropic (Claude): strongest on the hardest multi-file engineering and top of human-preference rankings; frontier output priced at or above Sol.
- Google (Gemini 3.1 Pro): the value champion — roughly $2 / $12 per 1M tokens and the widest free tier.
- OpenAI (GPT-5.6): a three-tier ladder plus Ultra multi-agent and the merged Codex harness, betting on agentic/coding workflows and cost-per-result rather than any single-benchmark crown.
Bottom line
The durable story is not one benchmark. It is that a frontier release now ships as a priced ladder (Sol/Terra/Luna), that the moat is moving to the agentic harness (Ultra, Codex, programmatic tool calling) rather than the raw checkpoint, and that who is allowed to ship, and to whom, is now a government-shaped variable. For anyone routing work across models by cost and capability, GPT-5.6 widens the menu and lowers the frontier’s floor — while leaving the vendor-vs-independent benchmark gap wide open.
Sources
- Launch, models, pricing, specs, Ultra: OpenAI — GPT-5.6 (403 on fetch; corroborated below), explainX, eden AI, QCode
- Independent analysis + benchmark caveats: Simon Willison, o-mega
- Restricted preview → GA / government request: TechCrunch, TechRadar, Medium — “Thirteen Days”
- Usage terms + temporary limit relaxation: OpenAI Usage Policies, BleepingComputer, Android Authority
- Competitive landscape / pricing comparison: tech-insider, LM Council benchmarks
- ITmedia (source the CEO flagged): OpenAI「GPT-5.6」一般提供 — 3モデルの価格と新たな利用条件