Anthropic’s Fable 5 came back from the dead this week, and nobody can agree on whether it’s broken or fine. OpenAI quietly shipped a custom inference chip with Broadcom. Meta told its staff they’ve caught GPT-5.5. The AI frontier isn’t moving in one direction anymore. It’s splitting into competing bets on infrastructure, safety, and cost, and this week made that painfully obvious.
Fable 5 Returns, But the Router Won’t Let Go
Claude Fable 5 went dark on June 12 after the U.S. Commerce Department slapped export controls on it, following an Amazon research report showing the model could identify software vulnerabilities and even produce working exploit code. Anthropic had about 90 minutes to comply, so they pulled it for everyone. Three weeks later, on July 1, it came back.
The comeback came with a catch. Anthropic trained a new safety classifier that blocks the specific hacking prompt from the Amazon report in over 99% of attempts, rerouting flagged requests to the weaker Opus 4.8. The problem: that classifier can’t tell the difference between a malicious exploit request and a routine debugging task. BridgeBench reported a collapse in coding scores, with debugging falling from 86.2 to 25.9 and refactoring dropping from 73.6 to 38.4. But here’s the twist: only 3 of 12 TypeScript debugging tasks actually reached Fable 5. The other 9 were intercepted by the classifier and sent to Opus 4.8, scored as zeros.
Arena.AI’s blind human-preference data told a different story. Document performance rose 34 points, expert text gained 25, creative writing ticked up 9. Frontend code slipped slightly but stayed within confidence intervals. So Fable 5 still performs like Fable 5 when prompts actually reach it. The issue isn’t model decay. It’s a paranoid router that treats “fix this vulnerability” as an attack.
Anthropic acknowledged the false positives and said they’ll refine the system, but gave no target date. They also proposed an industry-wide framework for scoring jailbreak severity, with Amazon, Microsoft, Google, and other Glasswing partners. Meanwhile, Mythos 5 (the same underlying model with fewer guardrails) stays locked to roughly 100 vetted U.S. organizations. The full Fable 5 redeployment post is at anthropic.com.
CrowdStrike Warns Mythos Could Compress Zero-Day Windows
CrowdStrike President Michael Sentonas told Observer that interest in Claude Mythos jumped after the company’s role in Project Glasswing became public. “The phones went crazy,” he said. Companies wanted access CrowdStrike couldn’t provide. They’ve been advising organizations on how to think about it ever since.
The core warning: AI could soon help discover and exploit vulnerabilities far faster than organizations can patch them. “Imagine a world where 200 vulnerabilities are discovered every day,” Sentonas said. “Now imagine those vulnerabilities being exploited almost immediately.” At $50 per million output tokens, Mythos isn’t cheap, but CrowdStrike hasn’t seen major internal surprises because they already use models like Opus for vulnerability scanning. The real fear is broader access. If attackers get comparable tools, defenders lose the race.
OpenAI’s Jalapeño Chip and GPT-5.6 Sol Preview
OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first custom inference accelerator. It’s a blank-slate design built specifically for LLM inference, not a repurposed general-purpose chip. Early testing shows performance per watt substantially better than current state-of-the-art. Engineering samples are already running GPT-5.3-Codex-Spark in the lab at production target frequency.
The chip went from design to tape-out in nine months, which OpenAI claims is the fastest ASIC development cycle ever in high-performance semiconductors. They used their own AI models to accelerate parts of the design process. That’s not a demo. That’s AI building the hardware that runs AI. Broadcom handles silicon implementation and networking, Celestica does board and rack systems. First deployment is planned for gigawatt-scale data centers by end of 2026.
Alongside the chip news, OpenAI previewed GPT-5.6 Sol on June 26, their strongest model yet. It introduces a new “ultra mode” that uses subagents to accelerate complex work, and sets a new state of the art on Terminal-Bench 2.1 for command-line workflows. On the cybersecurity side, GPT-5.6 Sol is competitive with Anthropic’s Mythos Preview on ExploitBench while using roughly one-third the output tokens. The catch: it’s in limited preview to a small group of trusted partners, with government coordination. OpenAI explicitly said they don’t believe this kind of government gating should become the default, but framed it as a short-term step toward broader availability.
OpenAI also shipped GeneBench-Pro on June 30, a research-level benchmark for computational biology with 129 problems across 10 domains. GPT-5.6 Sol scores 28.7% at the highest reasoning level. That’s up from below 5% when they started building the original GeneBench with GPT-5. They estimate the benchmark could be saturated by end of year.
Meta’s Watermelon Model Catches GPT-5.5
Alexandr Wang, Meta’s superintelligence chief, told staff at an internal town hall that Meta’s next model, codenamed Watermelon, has matched OpenAI’s GPT-5.5 on closely watched benchmarks. The model is still in training and uses an order of magnitude more compute than Avocado (Meta’s internal name for Muse Spark, released in April).
If accurate, this would be the clearest sign yet that Zuckerberg’s massive AI spending is producing results. Meta has trailed OpenAI, Google, and Anthropic in frontier AI for a while. The company now expects to spend $125-145 billion on infrastructure this year, up from the previous $115-135 billion forecast. Wang also said on X that an update to Muse Spark is coming soon with stronger coding and agentic abilities, and when asked when Meta would have a coding model on par with Claude Opus, he replied “pretty soon.” Meta declined to comment.
GLM-5.2: China’s Open-Weight Pressure Play
Z.ai released GLM-5.2, an open-weight model from Beijing that developers say approaches U.S. frontier systems at a fraction of the cost. Reuters reported rising use on platforms like OpenRouter, where pricing data shows it costs far less per token than leading U.S. models. Some executives are calling it a “mini DeepSeek moment.”
The model has drawn praise for coding, reasoning, and agentic AI. Unlike earlier Chinese models dismissed as budget tools, GLM-5.2 competes on quality. The advantage is strongest for startups and emerging markets where near-frontier performance at one-sixth the cost changes purchasing decisions. Western enterprise adoption in finance, healthcare, and defense will be slower due to security and compliance concerns, but the pricing pressure on OpenAI and Anthropic is real.
Mistral OCR 4 and Leanstral 1.5
Mistral shipped two notable releases. OCR 4 is their state-of-the-art document intelligence model, supporting 170 languages and returning bounding boxes, block classification, and inline confidence scores alongside extracted text. Independent annotators preferred OCR 4 over every leading OCR system tested, with win rates averaging 72%. It’s priced at $4 per 1,000 pages via API, runs in a single container for self-hosted deployments, and integrates with Mistral’s Search Toolkit for RAG pipelines.
Leanstral 1.5 is a free Apache-2.0 licensed formal verification model with 6B active parameters. It saturates miniF2F completely (100%), solves 587 of 672 PutnamBench problems, and hits state-of-the-art on FATE-H (87%) and FATE-X (34%). It also found 5 previously unknown bugs across 57 open-source repositories tested. The model is fully open-sourced on Hugging Face with a free API. Formal verification just got practical.
Quick Hits
Hugging Face – Hugging Face and Cerebras demonstrated a real-time speech-to-speech pipeline pairing Gemma 4 31B with Cerebras inference. The open, modular stack uses Parakeet for speech recognition, Gemma 4 for language, and Qwen3TTS for output. It already powers over 9,000 Reachy Mini robots in the wild. Full post here.
Google DeepMind – June brought a steady stream: computer use in Gemini 3.5 Flash, DiffusionGemma (4x faster text generation), Gemma 4 12B as a unified encoder-free multimodal model, and a research partnership with A24. Nothing single-day headline-grabbing, but solid infrastructure and model work across the board.
OpenAI Engineering – OpenAI published a deep dive on tracking down an 18-year-old race condition in GNU libunwind that was causing crashes in their Rockset data infrastructure. The debugging process involved treating core dumps like epidemiology data. It’s a fantastic engineering writeup. Read it here.
Rundown for July 4, 2026. Sources: Yellow, Anthropic, OpenAI, Google DeepMind, Mistral AI, Hugging Face.