Friday felt like two different industries sharing a calendar. Gimlet Labs, a 30-person startup that’s barely two years old, just priced itself at $3 billion on the idea that AI workloads shouldn’t care what chip they run on. Meanwhile, Anthropic’s Claude formalized the proof of Fermat’s Last Theorem into Lean, and the mathematician leading the rival human effort downloaded the result, compiled it, and conceded. One story is about squeezing more intelligence from every watt. The other is about what these models can already do when you hand them the right tools. Both are worth your morning coffee.

Gimlet Labs Lands $3B Valuation for the Anti-One-Chip Bet

Gimlet Labs raised $300 million in a Series B led by Andreessen Horowitz, landing at a $3 billion valuation. Total funding now sits at $392 million, and the round brought in two strategically loaded new names: Arm Holdings and Microsoft’s M12 venture arm. Sapphire Ventures, Menlo Ventures and Factory also came back for more.

Here’s what they bought. Gimlet runs what it calls the first multi-silicon inference cloud: instead of locking a workload to one vendor’s GPUs, its software splits an inference job into phases and routes each phase to whatever silicon is fastest and cheapest at that moment, whether that’s an Nvidia GPU, an AMD chip, a CPU, or a purpose-built accelerator. The company claims up to 10x throughput gains within the same power envelope. CEO Zain Asgar, a former Nvidia GPU architect, founded the company in 2023 and spent two years in stealth before launching last October.

Arm’s check is the tell. Arm-based server chips are quietly gaining share in inference because they’re cheaper per token than GPU-only stacks, and Arm has an obvious interest in software that makes its architecture a default rather than an afterthought. Microsoft’s M12 points the same direction, toward Azure’s own compute-allocation conversations. The risk sitting under the valuation: Nvidia’s Dynamo and Google’s TPU stack already do pieces of this natively, and the more successful Gimlet gets, the more it invites the giants to build the orchestration layer themselves. Still, at $3 billion for roughly 30 people, investors are clearly betting the middleman layer survives. That’s not a demo bet. That’s an infrastructure bet.

Claude Formalizes Fermat’s Last Theorem, and the Rival Team’s Leader Certifies It

Anthropic published research this week showing Claude produced a complete, machine-checkable Lean formalization of Fermat’s Last Theorem. To be precise about what happened: Andrew Wiles proved the theorem in 1994. What Claude built is the formalization, the translation of that proof into the Lean proof assistant such that a computer can verify every step with zero unproven placeholders. The claim: about 11 days of work, roughly 6 billion output tokens, largely autonomous, orchestrated through an open-source tool called Prove2Me. Anthropic’s first attempt without it failed.

The best part of this story isn’t in Anthropic’s post. Kevin Buzzard, the Imperial College mathematician who has led the community’s multi-year effort to formalize FLT in Lean, downloaded the public repository himself, ran the kernel checker, and confirmed it’s sorry-free, depends on only Lean’s three standard axioms, and matches Mathlib’s own definition of the theorem. Then he published a post titled “FLT: Anthropic has beaten me to it.” That’s adversarial verification from the person with the most expertise and the most reason to find a flaw, and he found none.

Wiles gets to keep his theorem. But an AI system producing an artifact a top mathematician couldn’t fault, in a proof system where no hand-waving survives, is a different kind of milestone than a benchmark score. The interesting question now isn’t whether models can formalize known results. It’s what happens when the target isn’t 360-year-old solved problems.

Microsoft Puts Its Image Models Where Its Margins Are

Microsoft AI opened public preview of MAI-Image-2.6 in Microsoft Foundry on Friday and shipped a companion model, MAI-Image-2.6-Flash, for high-volume work. The flagship ranks No. 2 on Arena for both text-to-image and image editing as of September 4. Flash generates images 2.8x faster than GPT-Image-2-Medium with 72% better efficiency, per Microsoft’s numbers. Output pricing runs $38 per million image-output tokens for the flagship, $19 for Flash. Arena’s estimate works out to roughly 4 cents an image.

Pair this with Thursday’s MAI-Transcribe-2 launch ($0.10 per hour of audio, 5.2% word error rate across 60 languages, 10x faster than GPT-Transcribe) and the playbook is unmistakable. Mustafa Suleyman’s lab picks a modality, ranks near the top of the preference leaderboards, then undercuts the frontier labs on price and routes everything through Foundry. On the earnings call, Microsoft was explicit that MAI models exist partly as an internal option for workloads where paying another provider would compress margins. This is procurement strategy wearing a product launch. Developers win either way: two credible image models at half the output cost of the flagship, live in preview today.

Quick Hits

xAI – Grok Bot got a template marketplace Friday (69 public Bots from 43 creators as of launch day) plus an enterprise tier with audit controls. Its launch agent, Haggle Bot, claims $100K+ in procurement savings in week one, though the figures mix measured savings with annualized estimates and no vendor names were published. Grok Bot also landed on iPad and Android, and the price dropped from $300/month to inclusion with Cursor Pro at $20/month. All of this one day after a Memphis compute outage took Grok down for 3.5 hours Thursday morning. Bumpy week for a product expanding this fast. Elon Musk says corrective action is underway.

Google – Lyria 3.5, its best music generation model, landed in the Gemini app and API on September 4 with more expressive vocals and richer arrangements. Gemini Spark can now manage your Google Photos library too: editing, album curation, even turning concert flyers into calendar events. Rolling out over the next few weeks to AI Pro and Ultra subscribers in the US. Quiet, steady productization.

Hugging Face – Transformers 2.0 arrived this week with 200+ new pretrained models, 30 additional languages, and up to 40% faster inference on common workloads. A decent farewell release before the Nvidia acquisition closes.

Meta – Muse Spark 1.3 keeps rolling out to Facebook, Instagram and the Meta AI app. The weights decision is still pending. Could be strategy. Could just be slow.


Rundown for September 5, 2026. Sources: Yellow, Anthropic, Microsoft AI, xAI, Google Gemini, Hugging Face, Meta AI.