The day’s the day isn’t between two models. It’s between two chip roadmaps, and both sides published benchmark numbers on the same day. That’s rare, and it makes for the most honest hardware news we’ve had in a year.

The Efficiency War Goes Public

OpenAI published first results for Jalapeño, its custom inference chip built with Broadcom. The headline claim: a 700W ASIC that beats Nvidia’s GB300 on power efficiency and response speed, in OpenAI’s own tests. The flashiest number is a claimed 104.3x edge on a mixed input/output token benchmark, and yes, auditors are already poking at it, because it compares work per dollar, not raw speed. Strip that away and the real number is still impressive: all-in system power of 1.18kW for Jalapeño versus 2.55kW for the GB300, against a GB300 running a superseded part.

Sarah Friar’s essay the same day, “The full stack behind abundant intelligence,” frames the strategy: chips, compute, models, and products compounding toward more useful intelligence at lower cost. Jalapeño is the physical layer of that thesis, and it just went from rumor to published data.

Nvidia answered within hours. Vera Rubin NVL72 systems deliver up to 30x more work per megawatt than GB300 NVL72 on agentic workloads, measured on SemiAnalysis’s AgentX benchmark at 160 tokens per second per user. That’s a counterpunch aimed squarely at anyone planning a fleet of always-on agents.

Here’s what this tells us: agentic AI is now deciding hardware strategy. Nobody’s pitching raw flops anymore. The unit that matters is tokens per kilowatt-hour, because a thousand agents running overnight burn real power. Both sets of numbers are vendor-run, so treat them as marketing with engineering behind it. But the direction is unmistakable: the next fight is energy per token, and it started this week.

Apple Goes 2nm With M6, M5 Ultra Pushes Local AI

Apple slipped the M6, its first 2nm chip, into the Mac mini on August 25, with up to 32GB of memory and a September 22 launch. Next to it, the M5 Ultra in the new Mac Studio is Apple’s most powerful chip ever, a quad-die design delivering up to 4.5x the AI compute of the M3.

That on-device number is the real story. Apple is betting that agents run locally, privacy-first, without phoning a data center. Different battlefield from the rack wars, same premise: the model that runs on your hardware without a metered API bill wins the device.

NemoClaw: One Malicious Website Can Poison a Local Agent

Oasis Security found a weakness in Nvidia’s NemoClaw, the tool that runs OpenClaw agents against a local Ollama server. A malicious webpage can reach that local Ollama unauthenticated, take control of the model, and plant hidden instructions inside it. One visit, no credentials, and your agent is executing rules you never wrote.

The root cause is mundane, which is why it works: NemoClaw leaves Ollama exposed to the network with no auth change. Local model stacks are not sandboxes. If you run agents locally, this is the patch of the week.

Quick Hits

Anthropic – launched Wellbeing Research Grants on August 25 to fund better ways of evaluating AI’s impact on human wellbeing, and it admits wellbeing is the hardest outcome to score.

SpaceXAI – Grok Voice now resolves more than 15,000 Starlink support and sales calls a day and closes 3,000+ orders a week. Voice agents doing real revenue work, not demos.

Ox AI – a stealth 1M-context model is live on OpenRouter, free until August 27, posting an 80% score on DeepSWE. OpenRouter won’t name the builder and tokenizer forensics point to Zhipu GLM. A week of free frontier-grade coding is a statement about pricing power.

Intuit – its AI “Big Bets” grew 34% and now drive 30% of revenue, per its August 25 investor update. AI on the income statement is no longer theoretical.

Wrtn – the Korean AI companion startup raised $72M at a $722M valuation on August 25, a signal that companion bots are still getting funded.


Rundown for August 26, 2026. Sources: OpenAI, NVIDIA, Apple, Anthropic, The Hacker News, PCMag, Intuit, 933 The Drive, OpenRouter.