[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%
Anthropic releases Claude Opus 5.5 as default model; Anthropic and competitors cut pricing 40-50%, overshadowing OpenAI's GPT-6 efficiency gains.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Anthropic releases Claude Opus 5.5 as default model; Anthropic and competitors cut pricing 40-50%, overshadowing OpenAI's GPT-6 efficiency gains.
Anthropic released Claude Opus 5.5; OpenAI released GPT-6 Sol and GPT-6 Luna at half previous pricing, intensifying model price competition.
Anthropic says its new Claude Opus 5.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents. In an announcement on Tuesday, Anthropic says Opus 5.5 comes with improvements to certain risky behaviors, including attempts to escape the company's testing sandbox. It's the first model released by Anthropic after CEO Dario Amodei announced plans to "pace the frontier," or slow down AI development. In recent weeks, several AI companies, including Anthropic, Google, and OpenAI, have reported that their AI models escaped containment and hacked third-party companies dur...
It’s been a busy few months for AI hype. At the end of April, Anthropic claimed that its model Claude Mythos is better at finding software vulnerabilities than most security experts. Then we had the OpenAI–Hugging Face hacking incident, after which Anthropic (proudly) and Meta (reluctantly) disclosed similar incidents involving their models. This was followed…
Anecdote: large company relying heavily on Claude Code for artifact generation creates workflow dysfunction despite high throughput expectations.
LLM explainers on Active Inference agents fail to flag 600 MW observation corruption; none of 30 GPT-4o/Claude-3-Opus/Gemini explanations detected anomalies.
Anthropic adds AGENTS.md support to Claude Code v2.1.277 for project instruction customization via modular architecture.
Researchers used Claude to reach an OpenAI employee account and sensitive GitHub data.
The revamped projects feature in Claude Code allows users to run multiple agents under the same roof, with a shared memory, goals, and library of files and artifacts. Similar to Grok Bot and other tools that manage groups of AI agents, each project has "threads" running different tasks in parallel, with a "coordinator" directing everything: Under the hood, each thread is a Claude Code cloud session working on its own branch and copy of the repo. The coordinator keeps work organized, but if any threads work on the same code, the overlap is resolved as a merge conflict just like any other PR. E...
Anthropic launches Life Sciences Verification Program granting researchers access to Claude models with relaxed safeguards for biology applications.
Study of 67,200 responses from GPT, Claude, and Gemini across 112 languages shows geopolitical bias variance in LLM answers about Ukraine war.
Anthropic merges Claude Cowork and Claude Chat into unified interface with persistent agent capabilities across web, desktop, mobile for Pro/Max users.
Benchmark: six frontier models play log(N)-Questions game; Claude Opus 5 underperforms (28/68 wins), revealing communication asymmetry gaps.
Google is inviting third-party agents, including Claude and Open Claw, into Google Home. | Photo by Jennifer Pattison Tuohy / The Verge Google is opening up its smart home to AI agents, letting tools like Claude and Open Claw access and control your connected devices and analyze your home's data using the standardized Model Context Protocol. Google Home MCP is a new integration that lets third-party AI agents control and monitor your smart home and act on your behalf. It "allows any AI agents that support MCP, including Google Antigravity, Claude, Hermes or Open Claw, to securely work with al...
Google is launching early access to a new MCP server for Google Home, allowing AI agents like Claude, ChatGPT, and others to control connected devices, review camera summaries, and access smart home activity using natural language.
Anthropic is initially releasing these features to Pro and Max plan subscribers.
Claude is getting a pair of new tools today: Docs and Slides. They'll let you create documents and presentations through Claude chats, which you can export, edit, and share with other users. As part of the announcement, Anthropic is also simplifying how Claude chats work, merging regular chats and Cowork into "one Claude," with all of its AI productivity tools available from any chat. Artifacts and Claude Design capabilities will be available through the new single interface as well. According to Anthropic, Claude "can now figure out what a task needs, so what Cowork and Design can do is avai...
Tests whether Claude 3.5 Haiku's rhyme planning via newline-resident features generalizes across seven open-weights models and six cross-layer transcoders.
A new WhatsApp Business MCP server lets developers use AI coding agents like Claude, Cursor, Codex, and ChatGPT to handle setup, messaging templates, testing, and troubleshooting.
Boris Cherny outlines Anthropic's production guardrails for Claude-generated code: linting, testing, fuzzers, automated review, and refactoring.
Some dangerous biology looks much like legitimate research, complicating AI safeguards.
Datasette security patches (1.0a39, 0.65.4) address vulnerabilities in public/private table isolation, audited with Claude and GPT models.
Last month, a Claude user noticed his account was consuming tokens even though he wasn't working. Anthropic has since warned users about hackers.
Anthropic says power users are key to its business - it's prioritized them even when it means cutting off other popular applications, like OpenClaw. But some of these same customers say Anthropic misled them into believing they'd get more out of a top-tier pricing subscription than they did. In an expanded class-action lawsuit filed today, a group of Claude subscribers say the company deceptively advertised the limits of its Max subscription tier. The lawsuit was brought by attorneys Monica Vaca and Kati Daffan, who both formerly worked at the Federal Trade Commission under Lina Khan. It's a ...
API benchmark scores for ChatGPT, Claude, Gemini systematically diverge from chatbot interface behavior by 3.4pp average; questions reliability of model eval standards.
Simon Willison used Claude Fable 5.1 to build a WebAssembly-based video compressor tool via Claude Code.
Public-market scrutiny will intensify pressure on the Claude maker’s unusual attempt to balance profit and purpose.
Study of substrate blindness in AI agents: Claude Opus 5, GPT-5.6-Sol, Gemini 3.7 Flash code generation ignoring memory/compute constraints.
Simon Willison's August newsletter covers OpenAI security incidents, game-playing agents (Fable 5, Sol 5.6), and Claude auto mode with model releases roundup.