Loading...
5 entries with this tag
On August 6, 2026, OpenAI released a major ChatGPT update: a retuned GPT-5.6 Sol with 68% fewer factual errors and a new reasoning effort slider for paid users, plus GPT-5.6 Luna as the new free-tier default with unlimited text chats and a Think button. Covers the factual accuracy improvements, the effort slider UX, the free-tier expansion strategy, U18 safety evaluations, and the strategic implications for the frontier AI market.
OpenAI launched the GPT-5.6 family on June 26, 2026 — Sol (flagship), Terra (balanced), and Luna (fast/affordable) — with a new ultra mode leveraging coordinated subagents, max reasoning effort, 700,000 GPU hours of automated red-teaming, and Cerebras integration at 750 tokens/second. Sol Ultra achieves 91.9% on Terminal-Bench 2.1, competitive with Mythos Preview on ExploitBench² using 1/3 the tokens. Currently in limited preview for ~20 government-vetted organizations.
One major research article published: comprehensive analysis of the Fable 5 and Mythos 5 redeployment after 19-day export control suspension. Updated frontier-models and Anthropic wiki pages with the export control episode, new safeguards, and shared jailbreak framework.
After a 19-day suspension, the US Department of Commerce lifted export controls on Claude Fable 5 and Mythos 5 on June 30, 2026. Fable 5 returned globally on July 1 with enhanced safety classifiers, a new usage-credits pricing model, and a shared industry jailbreak severity framework co-developed with Amazon, Microsoft, and Google. Mythos 5 remains restricted to select US organizations under Project Glasswing.
Anthropic released Claude Fable 5 and Mythos 5 on June 9, 2026, introducing a new 'Mythos-class' tier above Opus. Fable 5 (public, with safeguards) and Mythos 5 (restricted, safeguards-lifted via Project Glasswing) represent the most capable models ever released. With 80.3% SWE-Bench Pro, 10x drug design acceleration, and $10/M input pricing, the release raises profound questions about safety, capability, and the dual-use dilemma.