Speed is the new frontier. OpenAI just made GPT-5.6 Sol run 14 times faster with Cerebras hardware, and Google slashed Gemini 3.7 Flash pricing by 50% while jumping from 49% to 65.3% on DeepSWE. Both moves say the same thing: the race isn’t about smarter models anymore, it’s about making intelligence cheap and fast enough to embed everywhere. And IBM, fresh off a 25% stock collapse, just bet its consulting future on GPT-5.6. Progress and desperation, side by side.
OpenAI Goes Ultrafast: 750 Tokens Per Second With Cerebras
OpenAI previewed Ultrafast mode on August 13, and the numbers are genuinely startling. GPT-5.6 Sol running up to 14 times faster than standard processing, generating up to 750 output tokens per second. That’s not a benchmark improvement. That’s a different category of product. The kind of speed that changes what AI can do in real-time incident response, financial trading analysis, customer support, and live commerce.
Here’s what makes this interesting: Cerebras is powering it, not NVIDIA. OpenAI is explicitly diversifying its inference infrastructure, and the partnership signals that wafer-scale chips can compete on frontier model serving, not just research demos. The preview is limited to select customers, but OpenAI is already testing it internally for incident response and research workflows. Their own teams describe tightening overnight batch experiments into interactive working sessions. That’s not a demo. That’s production behavior.
The catch: capacity is limited, and there’s no public timeline for broad access. But the signal is clear. When frontier intelligence runs at 750 tokens per second, the bottleneck moves from the model to everything around it, your tools, your data pipelines, your ability to keep up.
Google’s Gemini 3.7 Flash: Half the Price, Twice the Coding Benchmarks
Google launched Gemini 3.7 Flash on August 13, just three weeks after 3.6 Flash. That cadence is aggressive even by Google’s standards. The improvements are concrete: DeepSWE v1.1 jumped from roughly 49% to 65.3%, FrontierCode 1.1 Main rose from 34.4% to 43.6%, and WebDev Arena Elo climbed from 1538 to 1588. This is a coding model that’s getting measurably better every three weeks.
The pricing is where it gets spicy. Introductory rates of $0.75 per 1M input tokens and $3.75 per 1M output tokens through the end of 2026. That’s half what 3.6 Flash launched at three weeks ago. Google is clearly willing to subsidize adoption, and the message to developers is blunt: switch now, save money, get better results. Outside testing from Harvey and Nunu.ai confirms the model punches above its weight class, landing near mid-sized frontier models at about half the cost.
This is an infrastructure play, not a feature play. Google isn’t competing on benchmarks alone. It’s competing on unit economics. If your agentic workflow runs hundreds of planning steps, tool calls, and retries, the per-token cost matters more than any single benchmark score. Google knows this. The 50% price cut is aimed directly at OpenAI’s enterprise base.
IBM Bets on GPT-5.6 After Worst Trading Day in Company History
IBM announced a deep partnership with OpenAI on August 13, folding GPT-5.6, Codex, and ChatGPT Work into its IBM Consulting Advantage platform. A dedicated OpenAI Practice staffed by thousands of consultants with expert-level certifications. IBM joins OpenAI’s Elite partner tier. The first industries in scope: financial services, government, telecom, and retail.
The timing is the story. This deal lands exactly one month after IBM shares fell 25% on July 14, the steepest single-day drop in the company’s trading history, eclipsing the October 1987 crash. CEO Arvind Krishna told investors IBM hadn’t adapted quickly enough as large deals failed to close. Q2 revenue hit $17.2 billion, up 1%, but adjusted earnings missed at $2.93 versus $3.01 expected. Management trimmed full-year revenue growth guidance to 4-5%.
So this is what a turnaround bet looks like. IBM is putting its consulting reputation on OpenAI’s models at the exact moment both companies need a win. Denise Dresser, OpenAI’s outgoing CRO (more on that below), framed it as combining transformation expertise with real business priorities. The security strand extends IBM’s earlier participation in OpenAI’s Daybreak Cyber Partner Program. Both companies plan to pair OpenAI models with IBM Autonomous Security for coordinated threat response. It’s either a smart hedge or a desperate move. Probably both.
OpenAI’s Daybreak Cyber Program Gets a Red Tier
OpenAI expanded its Daybreak cyber defense program on August 10 with two tiers. Daybreak Blue gives approved defenders access to GPT-5.6 Sol with safeguards tailored for defensive security work: vulnerability discovery, malware analysis, incident response, patch validation. Daybreak Red goes further, offering GPT-5.6-Cyber, a model trained specifically to reduce refusals on dual-use cyber tasks like exploit chain development and privilege escalation.
The numbers tell the story. GPT-5.6-Cyber completes 95% of advanced cybersecurity requests in OpenAI’s internal evaluation, compared to 1.5% for standard GPT-5.6 Sol and 2% with Daybreak Blue. The previous generation, GPT-5.5-Cyber, managed only 57.3%. Security researchers had been hitting persistent refusal walls. OpenAI clearly heard that feedback.
This is a careful balance. OpenAI is putting offensive-capable AI in trusted hands before threat actors deploy similar tools at scale. The program requires approval, and the access tiers create a gradient between general defensive work and legitimate offensive research. But the narrowing cyber defense window is real. If attackers get autonomous AI exploitation first, defenders need models that don’t refuse half their requests.
ChatGPT Gets Computer History and Ads Go Global
Two more OpenAI moves from this week deserve attention. Computer History launched in the ChatGPT Mac desktop app on August 13, rolling out globally to Pro, Business, and Enterprise users. It tracks activity across apps and websites, builds a timeline, and lets future ChatGPT sessions draw on that context. Think of it as persistent memory for your actual workflow, not just your chat history. EEA, UK, and Switzerland access is delayed, which tells you the privacy regulatory picture is still being navigated. Users must explicitly opt in under Settings, and no opt-in means no data collection.
Meanwhile, ChatGPT ads expanded to the UK, Mexico, Brazil, Japan, and South Korea on August 11. The ad pilot started in the U.S. in February, expanded to Canada, Australia, and New Zealand in March, and OpenAI reports no impact on consumer trust metrics with low dismissal rates. Free and Go tiers see ads. Plus, Pro, Business, and Enterprise don’t. Users can opt out of ads in the Free tier in exchange for fewer daily messages. That’s the trade: attention or money. Standard internet economics, now inside a chatbot.
OpenAI Appoints New CRO as Enterprise Gap Widens
Dali Rajic joins OpenAI as Chief Revenue Officer, replacing Denise Dresser (who notably appeared in the IBM partnership announcement, suggesting the transition is still in progress). Rajic comes from Wiz, recently acquired by Google, where he was President and COO. Before that, Zscaler and AppDynamics. This is a revenue operator who has scaled cybersecurity and enterprise SaaS companies through major growth phases.
The backdrop: OpenAI’s enterprise data shows frontier firms (top 10% of AI usage) now generate 8.3x as many output tokens per active user as typical firms, up from 2.6x in January. Codex generates 64% of combined Codex and ChatGPT output tokens among enterprise customers. Weekly active Codex users grew 108x in legal, 41x in sales, and 26x in marketing since February. The gap between companies getting real value from AI and those just dabbling is widening fast. Rajic’s job is to close that gap, or at least monetize it.
Quick Hits
Google DeepMind – Beyond Gemini 3.7 Flash, DeepMind posted two other August updates: a sign language AI model aimed at putting communication tools in users’ hands, and WeatherNext, an AI model that achieved a breakthrough in forecasting cyclones. Neither got the attention of the Flash launch, but both fit DeepMind’s pattern of applying frontier models to underserved domains.
Hugging Face – New blog post on August 13 covering Strands Agents, LeRobot, and Hugging Face Storage Buckets. The pitch: record, train, and deploy from one place. It’s an integration story, connecting robotics datasets (LeRobot), agent frameworks (Strands), and storage infrastructure into a single workflow. Practical, not flashy.
Mistral AI – Pushing hard on sovereign AI for Europe. The latest post covers in-region inference, open models, and new European infrastructure. Mistral is positioning itself as the EU’s domestic AI champion, and the message is aimed at governments and enterprises worried about US model dependency. Whether the economics work is a separate question from whether the political timing is right.
Rundown for August 15, 2026. Sources: Yellow, OpenAI, Google DeepMind, Hugging Face, Mistral AI.