Grok 4.7 Bets on Longer Coding Jobs xAI's new Grok 4.7 is built to grind through multi-hour tasks, check its own work, and keep the same price as the model before it.
Anthropic's Opus 5.5 Gets Cheaper and Smarter Anthropic's new Opus 5.5 beats its larger Fable model on many benchmarks while costing less to run, all under fresh safety limits.
Gemini 3.8 Live Learns to Talk Back Google's new voice models can reason while they speak, switch languages mid-sentence, and finish tasks in the background while the conversation keeps going.
Baking Safety Into Open AI Models Anyone can strip the guardrails off an open AI model. Baseten's new research arm wants safety built in from the start, not bolted on.
Plugin4Shell Bug Hit AI Coding Agents A flaw in Codex, Claude Code, Gemini CLI and Copilot let attackers swap a trusted plugin for malicious code, no developer click required.
DeepSeek's New Flash Model Goes Small DeepSeek says V4.1-Flash is the smallest model in a new architecture family, with vision built in from the start.
OpenAI's 10,000-Agent Math Claim OpenAI says a swarm of agents cracked a Navier-Stokes result, but the proof, the details, and the field's verdict are all still missing.
GPT-6 Astra Beats Portal by Itself OpenAI's new flagship model played through all of Valve's Portal on its own, and one hobbyist caught the whole 24-hour run.
OpenAI's GPT-6 Astra Hits a Cyber Red Line OpenAI says its new flagship crossed a "Critical" cybersecurity threshold, a first for the company and a puzzle for every enterprise using AI agents.
Apple Bets on Local AI With New Macs Apple's refreshed Mac mini and Mac Studio are pitched at developers who want to run AI models at home instead of renting cloud tokens.