The story today isn’t one model getting smarter. It’s many models getting combined, and the compute to run them running out. Nous Research dropped an open-source ensemble that beats any single frontier model. Google told Meta it can’t have all the Gemini it wants. And inside OpenAI, Codex agents now produce 99.8% of output tokens. The chatbot era at work is ending, and the infrastructure era is getting crowded.
Hermes MoA 2.0: Open-Source Ensemble Beats GPT, Claude, and DeepSeek Alone
Nous Research released Hermes Mixture of Agents 2.0 on Sunday, and it does something deceptively simple: it queries GPT, Claude, and DeepSeek in parallel, collects their outputs, and synthesizes a final response. The result outperforms each component model individually across reasoning, coding, and instruction-following benchmarks. The margin is biggest on long-horizon reasoning tests, exactly where single models tend to lose coherence.
The framework stays open-source. Researchers can inspect the architecture, swap base models, and adapt the ensemble for specific use cases. That matters because it lowers the barrier for teams that want top-tier reasoning without paying frontier API costs on every inference call. You still pay for the base model calls, but the orchestration layer is free.
Here’s what this tells us: model diversity, not a single dominant model, may define the next phase of AI deployment. Andrej Karpathy cautioned this week that agent-first development risks repeating earlier research mistakes. Nous Research takes a middle path, using strong foundation models as inputs while adding an orchestration layer on top. The approach hasn’t been tested against the absolute latest frontier releases like Claude Sonnet 5 or GPT-5.6 Sol, so the benchmark picture may shift. But the direction is clear. Ensemble architectures are a serious path to capability gains that individual training runs can’t easily replicate.
Google Rations Gemini to Meta as Compute Demand Outruns Supply
Google restricted Meta’s access to Gemini AI models around March, unable to supply the compute Meta wanted. The shortfall disrupted several internal Meta AI projects tied to coding, advertising tools, and content moderation across Facebook and Instagram. Managers told engineers to ration AI tokens more sparingly. In May, Google made the caps formal, imposing usage limits across Gemini apps. Access now scales with available capacity, not with how much a customer is willing to spend.
This is the part that should worry anyone building on outside AI platforms. A signed enterprise contract no longer guarantees the compute a company plans around, regardless of price. Google Cloud’s order backlog nearly doubled to $460 billion. CEO Sundar Pichai acknowledged the strain on the earnings call, saying the company was “compute-constrained in the near term.” Cloud revenue cleared $20 billion in a single quarter for the first time, up roughly 63% year over year. Demand isn’t the problem. Supply is.
The knock-on effects are real. Meta sped up its pivot to an in-house model called Muse Spark and plans up to $135 billion in AI spending this year. Google agreed to pay SpaceX roughly $920 million a month for about 110,000 Nvidia GPUs as a stopgap. For every dollar of committed demand, Google spends only about 40 cents on new capacity, so the gap keeps widening. This is infrastructure play, not feature play. The companies that own their compute will have a structural advantage the ones renting it won’t.
Codex Now Generates 99.8% of Output Tokens Inside OpenAI
A new study from OpenAI, Columbia Business School, Wharton, and Duke University finds that Codex agents now produce 99.8% of output tokens OpenAI employees generate. ChatGPT, the company’s flagship consumer product, has been overtaken internally by its own agent tool. Among outside organizations, Codex accounts for 63.3% of output tokens. Among individual users, just 16.5%. The split tells you where agentic AI is heading first: organizations with deep training and no cost limits.
The sharpest growth came from workers far outside engineering. Non-developer adoption climbed 137 times among individuals and 189 times at organizations since August 2025. Legal, finance, and recruiting teams at OpenAI crossed into majority Codex use by April 2026, months behind engineers but with much steeper adoption curves. Task complexity rose with that spread. The share of individuals handing over jobs estimated to need more than eight hours of human effort hit 25.6%, up from 2.1% in December 2025. More than one in ten users run three or more agents simultaneously each week.
Median output per OpenAI worker rose at least tenfold across roles since November 2025, reaching 13 times for lawyers and more than 50 times for researchers. The authors caution that OpenAI is an unusually easy home for agents, so its figures overstate the typical firm. Even so, the internal shift previews where wider adoption is going. Human value is moving toward setting tasks, checking output, and steering multiple agents at once. That’s not a demo. That’s production, at least inside one company.
Anthropic Tops OpenAI at $965B Valuation
Anthropic closed a $65B Series H round at a $965B post-money valuation, moving past OpenAI’s earlier private-market benchmark of around $730B. The round was led by Altimeter Capital, Dragoneer, Greenoaks, and Sequoia Capital. About $15B came from hyperscalers, including $5B from Amazon, which has tied Claude more closely to AWS. Micron Technology also participated, linking Anthropic’s funding to the memory and chip supply chain behind AI infrastructure.
Anthropic’s revenue run rate has crossed $47B, according to CFO Krishna Rao. That number suggests Claude is moving beyond pilots into daily workplace use. The valuation signals that private-market confidence can move quickly when enterprise demand and infrastructure backing align. For OpenAI, the bar just went up. Both companies are now racing toward the $1 trillion mark, and the contest for enterprise AI is tighter than it has ever been.
Palantir’s Karp Declares War on Token Pricing
Palantir CEO Alex Karp went on live television Wednesday and blasted the token-based pricing model behind OpenAI and Anthropic. He told viewers that companies pour money into tokens yet capture little real value, even as each new model costs more. “Something has gone completely wrong,” he said. He argued the arrangement lets labs pocket recurring fees while quietly absorbing a client’s proprietary data and competitive edge over time. Palantir stock rose nearly 8% the same session.
The comments landed days after Palantir widened its partnership with Nvidia, folding open Nemotron models into secure government and classified infrastructure deployments. Palantir also published a nine-point manifesto on data sovereignty, warning companies against handing strategic information to outside providers too cheaply. The pitch is clear: rivals sell access, Palantir sells control. Whether that message resonates beyond the government contracts where Palantir already dominates is an open question. But the frustration Karp voiced is real, and it’s shared by corporate buyers watching their AI bills climb without clear productivity gains.
Quick Hits
Anthropic – Fable 5 returned globally July 1, alongside a proposed industry-wide framework for scoring jailbreak severity with Amazon, Microsoft, Google, and other Glasswing partners. Safety coordination across labs is quietly becoming its own product category.
OpenAI – GeneBench-Pro launched as a research-level benchmark for computational biology, with GPT-5.6 Sol scoring 28.7% at the highest reasoning level. The company also previewed GPT-5.6 Sol with a new ultra mode using subagents, and unveiled the Jalapeno inference chip with Broadcom, designed from scratch for LLM inference with a nine-month tape-out cycle.
Mistral AI – OCR 4 shipped with bounding boxes, block classification, and inline confidence scores across 170 languages. Leanstral 1.5 saturated miniF2F completely and found 5 previously unknown bugs across 57 open-source repositories. Vibe agent launched as a unified work and code agent with VS Code extension.
Hugging Face – Kernels project got a major redesign with a new repository type, trusted publishers, and code signing via Sigstore. Partnered with Cerebras to bring Gemma 4 to real-time voice AI with a speech-to-speech pipeline powering 9,000+ Reachy Mini robots. Every Eval Ever integration now cross-posts evaluation results to model pages.
Google DeepMind – June batch still the latest: Gemini Omni, Nano Banana 2 Lite, computer use in Gemini 3.5 Flash, DiffusionGemma (4x faster text generation), and Gemma 4 12B unified multimodal model. No new posts since the June cluster.
Rundown for July 6, 2026. Sources: Yellow, Anthropic, OpenAI, Mistral AI, Hugging Face, Google DeepMind.