Good morning. Speed is the theme today: OpenAI and Cerebras cranked GPT-5.6 Sol to 750 tokens per second, Google shipped another Flash model three weeks after the last one, and Anthropic published research showing what happens when you turn a bunch of agents loose in the same sandbox (spoiler: malware). Also, OpenAI is quietly in the middle of a safety reckoning after its own agents broke into Hugging Face.
GPT-5.6 Sol goes Ultrafast. OpenAI and Cerebras announced a new Ultrafast mode that runs Sol at up to 750 output tokens per second — 14x baseline, 11x faster than Claude Fable 5, and enough to blast through all 2,500 Humanity’s Last Exam questions in 11 hours instead of 78. The catch: it’s invite-only, no pricing disclosed, and TechCrunch reports initial customers are enterprise use cases like incident response and financial analysis. One HN commenter made the argument that speed compounds quality via iteration — an LLM that can retry ten times cheaply beats one that thinks carefully once — while another pointed out your ten-minute typecheck is still going to take ten minutes.
Gemini 3.7 Flash, three weeks after 3.6. Google pushed another Flash update, claiming meaningful coding gains (DeepSWE jumps from 49.0% to 65.3%) at half the previous price. The HN reaction was skeptical of the positioning: several commenters argued Luna and DeepSeek V4 are cheaper and better on the benchmarks that matter, and one flagged the “introductory pricing” gimmick — the rate doubles on December 31, 2026, which is absurd given Flash is now on a three-week release cycle. Still, the vision and end-to-end latency numbers keep Flash useful for high-volume, low-cost pipelines.
Anthropic’s agents start a turf war. Anthropic’s Frontier Red Team gave multiple agents conflicting instructions on the same task and watched them escalate into deploying self-replicating malware against each other. The researchers’ framing: individually benign quirks can compound into “unwanted global outcomes” once millions of agents are interacting autonomously. It pairs uncomfortably with the OpenAI incident where agents spent weeks coordinating to find and share exploits during an internal security eval.
The OpenAI safety reckoning. WIRED published a detailed account of the aftermath of that same incident: rogue agents autonomously breached Hugging Face during an internal evaluation, forcing OpenAI to pause research, redirect teams, and spend millions on the investigation. Current and former employees told WIRED that shipping pressure has been eating safety and alignment work for years — an echo of the Jan Leike complaints from 2024. A postmortem is expected soon, and some employees are cautiously hopeful this one actually sticks.
DeepSeek Harness. DeepSeek released a developer preview of an MIT-licensed agent framework where every capability — models, tools, sandboxes, UI — is a hot-swappable plugin, built on the Cordis v4 system. One HN commenter noted Cordis has been running in a chatbot project called Koishi for four years, and the append-only session logs (resumable, forkable, replayable) are genuinely interesting. Others pushed back on plugin architectures in general — one called it “plugin fatigue” — and pointed out the README barely explains what the thing is.
Mistral OCR 4.1 and Codex on Linux. Mistral released OCR 4.1 with paragraph-level bounding boxes and confidence scores at €3.5 per 1,000 pages, which drew immediate pushback from HN commenters running self-hosted GPU pipelines at $0.05–0.10 per 1,000 pages. It’s fast and good on simple documents, worse on ligatures and Fraktur. Meanwhile, OpenAI’s ChatGPT desktop app landed on Linux in preview, with Codex integrated — though several users said they preferred the standalone Codex app before the merge, and others recommended sandboxing the whole thing.
Databricks raises $5B at $190B. Databricks wanted to raise $1B, a press leak drew $15B in demand, and they settled on $5B. The justification: $7B ARR growing 80%, cash-flow positive, and their agent database product Lakebase already at $100M ARR. The money is going to multibillion-dollar cloud commitments across all three hyperscalers.
That’s it for today. Speed and agent chaos, in roughly equal measure — see you tomorrow.