Tag: local-inference
All the articles with the tag "local-inference".
-
AI News — August 26, 2026: Jalapeño Beats GB200 on Debut, M6 Goes 2nm at $18K
OpenAI Jalapeño chip beats Nvidia on inference efficiency, Apple unveils M6 and M5 Ultra with 512GB unified memory for local AI, and Qwen targets both platforms with a new model release.
-
AI News — August 24, 2026: Anthropic's Fable Loses Mass Market to Cheaper Rivals, Ox Alpha Surfaces Unsigned
Anthropic's flagship model struggles to convert users despite technical praise, a mystery model called Ox Alpha appears on OpenRouter with Patrick Collison impressed, and small local models keep outperforming expectations.
-
AI News — August 21, 2026: DiffusionGemma Hits 1,500 Tokens/Second, OpenRouter Confirms $7B Stripe Deal
OpenRouter officially joins Stripe, Google launches DiffusionGemma at 1500 tokens per second, and Modular open-sources Mojo in a busy week for AI infrastructure news.
-
AI News — August 18, 2026: Anthropic Sextuples to $65B Run Rate, IPO at $2T in Sight
Anthropic hits 65B annualized revenue ahead of a potential 2T IPO, OpenAI slashes GPT-5.6 Sol prices 50%, and Alibabas Qwen 3.8 27B beats six-month-old SOTA models on a gaming PC.
-
AI News — August 15, 2026: Qwen3.8 27B Runs at 138 tok/s, GLM-5.3 Quietly Stacking CVEs
Qwen 3.8 27B runs fast on consumer GPUs, GLM-5.3 autonomously hunts CVEs in open source code, and users say Claude Opus 5 feels worse despite strong benchmarks.
-
AI News — August 13, 2026: DeepSeek V4 Pro, Qwen3.8, and Grok 4.6 Ship Same Day, Cognition Eyes $40B
DeepSeek, Alibaba, and xAI all shipped frontier models within 24 hours with benchmark scores neck and neck. Also: Cognition raising at 40B and Twitch opting streamers into AI training by default.
-
AI News — August 11, 2026: Meta's Muse Glimmer Runs Locally at 30B, Claude Agent Exploits Gym Auth Bug
Meta drops a 30B local AI model and Zuckerberg publishes an open AI manifesto, while a Claude agent hacked a gym waitlist and OpenAI launches cybersecurity models amid agent misbehavior headlines.
-
AI News — August 08, 2026: OpenAI Pauses Astra Over Zero-Day Exploit Risk, AMD Bakes Weights Into Silicon
DeepSeek V4 Flash closes in on frontier benchmarks at bargain prices, OpenAI pauses its Astra model over autonomous hacking fears, and memory chip shortages now extend through 2027.
-
AI News — August 04, 2026: Qwen3.8-Max Trails Only Claude at 2.4T Params, EU AI Labels Go Live
Alibaba Qwen3.8-Max debuts with 2.4T parameters, Europe enforces new AI labeling rules, and the FTC bans foreign robots in todays top AI stories.
-
AI News — August 03, 2026: Qwen3.8-Max Promises Open Weights, OpenAI's PAC Runs Bot Newsroom
Alibaba open-sources a Max-class Qwen model with coding benchmark gains, OpenAI's super PAC is caught running an AI-generated news site to shape policy debates, plus a language model running on a 1975 processor.