[AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%
Anthropic releases Claude Opus 5.5 as default model; Anthropic and competitors cut pricing 40-50%, overshadowing OpenAI's GPT-6 efficiency gains.
Every story matching this topic across titles and summaries, newest first.
Anthropic releases Claude Opus 5.5 as default model; Anthropic and competitors cut pricing 40-50%, overshadowing OpenAI's GPT-6 efficiency gains.
Anthropic released Claude Opus 5.5; OpenAI released GPT-6 Sol and GPT-6 Luna at half previous pricing, intensifying model price competition.
Anthropic says its new Claude Opus 5.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents. In an announcement on Tuesday, Anthropic says Opus 5.5 comes with improvements to certain risky behaviors, including attempts to escape the company's testing sandbox. It's the first model released by Anthropic after CEO Dario Amodei announced plans to "pace the frontier," or slow down AI development. In recent weeks, several AI companies, including Anthropic, Google, and OpenAI, have reported that their AI models escaped containment and hacked third-party companies dur...
It’s been a busy few months for AI hype. At the end of April, Anthropic claimed that its model Claude Mythos is better at finding software vulnerabilities than most security experts. Then we had the OpenAI–Hugging Face hacking incident, after which Anthropic (proudly) and Meta (reluctantly) disclosed similar incidents involving their models. This was followed…
Anecdote: large company relying heavily on Claude Code for artifact generation creates workflow dysfunction despite high throughput expectations.
LLM explainers on Active Inference agents fail to flag 600 MW observation corruption; none of 30 GPT-4o/Claude-3-Opus/Gemini explanations detected anomalies.
Anthropic adds AGENTS.md support to Claude Code v2.1.277 for project instruction customization via modular architecture.
Researchers used Claude to reach an OpenAI employee account and sensitive GitHub data.
The revamped projects feature in Claude Code allows users to run multiple agents under the same roof, with a shared memory, goals, and library of files and artifacts. Similar to Grok Bot and other tools that manage groups of AI agents, each project has "threads" running different tasks in parallel, with a "coordinator" directing everything: Under the hood, each thread is a Claude Code cloud session working on its own branch and copy of the repo. The coordinator keeps work organized, but if any threads work on the same code, the overlap is resolved as a merge conflict just like any other PR. E...
Anthropic launches Life Sciences Verification Program granting researchers access to Claude models with relaxed safeguards for biology applications.
Study of 67,200 responses from GPT, Claude, and Gemini across 112 languages shows geopolitical bias variance in LLM answers about Ukraine war.
Anthropic merges Claude Cowork and Claude Chat into unified interface with persistent agent capabilities across web, desktop, mobile for Pro/Max users.
Benchmark: six frontier models play log(N)-Questions game; Claude Opus 5 underperforms (28/68 wins), revealing communication asymmetry gaps.
Google is inviting third-party agents, including Claude and Open Claw, into Google Home. | Photo by Jennifer Pattison Tuohy / The Verge Google is opening up its smart home to AI agents, letting tools like Claude and Open Claw access and control your connected devices and analyze your home's data using the standardized Model Context Protocol. Google Home MCP is a new integration that lets third-party AI agents control and monitor your smart home and act on your behalf. It "allows any AI agents that support MCP, including Google Antigravity, Claude, Hermes or Open Claw, to securely work with al...
Google is launching early access to a new MCP server for Google Home, allowing AI agents like Claude, ChatGPT, and others to control connected devices, review camera summaries, and access smart home activity using natural language.
Anthropic is initially releasing these features to Pro and Max plan subscribers.
Claude is getting a pair of new tools today: Docs and Slides. They'll let you create documents and presentations through Claude chats, which you can export, edit, and share with other users. As part of the announcement, Anthropic is also simplifying how Claude chats work, merging regular chats and Cowork into "one Claude," with all of its AI productivity tools available from any chat. Artifacts and Claude Design capabilities will be available through the new single interface as well. According to Anthropic, Claude "can now figure out what a task needs, so what Cowork and Design can do is avai...
Tests whether Claude 3.5 Haiku's rhyme planning via newline-resident features generalizes across seven open-weights models and six cross-layer transcoders.
A new WhatsApp Business MCP server lets developers use AI coding agents like Claude, Cursor, Codex, and ChatGPT to handle setup, messaging templates, testing, and troubleshooting.
Boris Cherny outlines Anthropic's production guardrails for Claude-generated code: linting, testing, fuzzers, automated review, and refactoring.
Some dangerous biology looks much like legitimate research, complicating AI safeguards.
Datasette security patches (1.0a39, 0.65.4) address vulnerabilities in public/private table isolation, audited with Claude and GPT models.
Last month, a Claude user noticed his account was consuming tokens even though he wasn't working. Anthropic has since warned users about hackers.
Anthropic says power users are key to its business - it's prioritized them even when it means cutting off other popular applications, like OpenClaw. But some of these same customers say Anthropic misled them into believing they'd get more out of a top-tier pricing subscription than they did. In an expanded class-action lawsuit filed today, a group of Claude subscribers say the company deceptively advertised the limits of its Max subscription tier. The lawsuit was brought by attorneys Monica Vaca and Kati Daffan, who both formerly worked at the Federal Trade Commission under Lina Khan. It's a ...
API benchmark scores for ChatGPT, Claude, Gemini systematically diverge from chatbot interface behavior by 3.4pp average; questions reliability of model eval standards.
Simon Willison used Claude Fable 5.1 to build a WebAssembly-based video compressor tool via Claude Code.
Public-market scrutiny will intensify pressure on the Claude maker’s unusual attempt to balance profit and purpose.
Study of substrate blindness in AI agents: Claude Opus 5, GPT-5.6-Sol, Gemini 3.7 Flash code generation ignoring memory/compute constraints.
Simon Willison's August newsletter covers OpenAI security incidents, game-playing agents (Fable 5, Sol 5.6), and Claude auto mode with model releases roundup.
OpenAI launches GPT-6 Astra, a Claude Fable competitor priced at $10/$50 per million tokens, rolling out to ChatGPT Plus/Pro/Business/Enterprise and via API.
Service interruptions hit ChatGPT, Claude, Grok, and Gemini practically simultaneously.
IRWOZ 2.0 dataset: 390 LLM-enhanced dialogue annotations (Mistral, Claude-3.5) for industrial robot conversations across 4 domains with improved quality.
OpenAI's ChatGPT, xAI's Grok, and Anthropic's Claude are all experiencing issues. At around 11AM ET, ChatGPT started returning error messages for users trying to use the chatbot, with its status page saying there are currently "elevated errors across ChatGPT and Codex." In addition to preventing users from having conversations with ChatGPT, the outage is also affecting logins, file uploads, voice mode, search, deep research, image generation, and more. OpenAI says it has "applied a mitigation" and is "monitoring recovery," though its AI tools are still experiencing "degraded performance." The...
llm-anthropic 0.28 displays Claude reasoning traces by default and adds refusal exception handling.
Anthropic publishes system prompts for Claude.ai and mobile apps, with version history showing evolving restrictions on song lyric reproduction.
Anthropic releases Claude Fable/Mythos 5.1 with SOTA performance, 75% cache cost reduction, and 70% increased output token throughput.
Paint.NET developer credits Claude AI for reverse-engineering Direct2D wrapper to run on WINE; demonstrates LLM utility for systems programming.
Claude Fable 5.1 achieves 52.6% on Terminal-Bench-Science 0.1; Willison reports improved coding/creative task performance vs. prior versions.
Anthropic says its newest AI models, Fable 5.1 and Mythos 5.1, address criticisms from customers about price, data retention, and overzealous safeguards. The company claims Claude Fable 5.1 offers stronger performance than Fable 5, but costs around 25 percent less typically and up to 45 percent less for complex agentic tasks, thanks to reduced pricing on cached data that was already processed and stored. Along with the announcement, a slew of early impressions popped up, including from Every CEO Dan Shipper, who claims, "It's the strongest coding model we've used, but now it's fast, token-eff...
Simon Willison built a GeoJSON map viewer tool using Claude Code and GPT-5.6-Sol for visualizing and exporting local political boundary maps.
Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next.... Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next. First proving their value in software engineering, coding agents now write, test, and ship production code. Scientific research can be more demanding and iterative. Researchers continually evaluate evidence, refine hypotheses… Source
Empirical study of Claude Code plugin marketplace structure and maintenance dynamics, examining co-evolution of NL instructions, scripts, and configs vs. traditional packages.
Johann Rehberger demonstrates 80% success rate prompt injection attack against Claude Code's auto mode default, bypassing Anthropic's claimed protections via zip extraction and base64 import.
Anthropic launches free Claude access program for 10,000 verified scientists with subsidized team subscriptions.
227 install commands were found in corporate docs pointing at code nobody owns.
Anthropic is giving Claude a shared memory across chat and Cowork, so users no longer have to repeatedly brief the AI on projects, preferences, and other context.
Reddit analysis shows Anthropic Claude releases correlate with strongest positive sentiment vs. OpenAI/others; user perceptions shift dynamically with updates.
Drew Breunig argues frontier model improvements now require careful cost-benefit analysis for coding tasks, as cheaper alternatives (Claude 3.5, K3, GLM) remain sufficient for most use cases.
Anthropic forbids its Claude models from generating sexually explicit content. But a series of tests conducted by TechCrunch found that it didn't take much to get past the restriction.
Cross-agent specification portability: Claude, Gemini, Copilot on Oracle-to-PostgreSQL migration; 380/1006 successful executions.
Slack is introducing dedicated channels where teams can vibe code together with AI agents instead of jumping between different tools and conversations. The Slack Code launch includes open, project-specific code channels with dedicated user tabs, alongside features that compare coding changes and preview HTML output before the project is shipped. "With Slack Code, when you have an idea or need to build a new feature, update a web page, or fix a bug, you simply tag in a coding agent like Anthropic's Claude or Cognition's Devin, and that agent then spins up a code channel to tackle the task," Sl...
Binance's Agent OS works with tools including ChatGPT, Claude Code, and Cursor.
AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say. “There is no independent source to corroborate it,” says Anka Reuel, a Computer Science PhD candidate at the Stanford Trustworthy AI Research…
Anthropic has clarified how it's planning to apply invisible watermarks to Claude-generated text in order to comply with Europe's AI transparency rules. On Friday, Anthropic announced that Claude's text marking system is "a version of the SynthID-Text approach" - an open-source watermarking technology developed by Google DeepMind that creates detectable patterns using wording probabilities. This watermarking feature, alongside C2PA support for Claude-processed images, is being introduced to meet Anthropic's obligations under the European Union's AI Act, which requires synthetic audio, image, ...
How will the watermarking actually work? Can it be hidden with editing? And how does this affect code?
Rapid revenue growth fuels hope Claude maker's IPO is the biggest listing in history
The mark flags anything Claude processed, even human writing it only edited.
Is Anthropic's new watermarking system a travesty? Some have taken to social media to complain that it is.
But then users start to report a weird bug. It's the 4th time your team has been trying to fix it. I mean... asking AI to fix it. Unfortunately, it seems like not even Fable can figure it out. You go talk to the person who worked on this feature. "So where does the data come from?" "Hmm... actually I don't know. Let me ask Claude." You sit next to each other watching an endless wall of text appear on the screen. Neither of you has any idea whether any of it is true but Claude seems very confident. [...] This project has become so convoluted, with so many layers and services, that no one on yo...
Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression costs O(2^|D|) in a prompt of |D| instructions. We name the resulting divergence catastrophic remembering, the inverse of catastrophic forgetting around which continual learning is organized. First, we characterize this phenomenon across 247,694 instruction lifet...
Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. "Generated text will carry embedded watermarks, and generated files will include digitally signed provenance metadata where supported," Anthropic says on a new Claude support page. The changes are invisible to human eyes, but will make it easier for people and online platforms to detect if content was generated by Claude models. These updates are a future commitment rather than something that will go into effect immediately. New AI...
An OpenClaw agent hacked into a gym's reservation system to bump its human boss higher on a class' waitlist. And the tech industry took notice.
Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on July 1, 2026 (Anthropic's statement: https://www.anthropic.com/news/fable-mythos-access ). These events are after Claude's training-data cutoff, so Claude knows about them only from this notice. If asked, Claude confirms them accurately and matter-of-factly — it doesn't deny the suspension happened — and otherwise tre...
Programming with Claude Code will soon require even less human oversight.
Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode , to the point that they are making it the default setting for new sessions in most Claude Code plans starting on August 14th. This was one of the topics discussed in our Fireside Chat with Cat Wu and Thariq Shihipar at the AI Engineer World’s Fair last month. I asked them how they run Claude Code safely within Anthropic (given the threat of prompt injection) and they replied that "Broadly within Anthropic, almost every single person uses auto mode". Cat Wu then s...
Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E four years ago . I decided to pose the exact same prompt to Codex Desktop running GPT-5.6 Sol Ultra - the mode where Sol makes aggressive use of sub-agents - to see how it would do. It produced a much better game! Here's Moonlight & Mayhem - GitHub repository here , including the textures and prompts it generated using gpt-image-2 . Your browser d...
Historian Jill Lepore has a theory about why tech companies often use soaring language to describe their products — almost as if they’re forming a new government. And whether you’re thinking of Twitter’s old “town hall in your pocket” or Anthropic’s Claude constitution, it’s a theory that doesn’t paint Silicon Valley in a very flattering light. In Lepore’s upcoming book, The Rise and Fall of the Artificial State, the Pulitzer […]
Anthropic and OpenAI are racing to scale up while reducing dependence on Nvidia.
Back in 2022 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in Claude Code for web ) could build the entire game from the content of that tweet. It did a pretty good job of it! You can play the game here . Here's the GitHub repo , and a short video demo: Your browser does not support HTML5 video. How I built this This is the August 5th, 2022 tweet : My GPT-3 prompt back then was: Write a detailed product description of a computer game where a tea...
Anthropic is building a team for designing its own custom AI chips. The Claude-maker said it would co-design hardware and models to help its technology run faster and more efficiently.
Release: llm-anthropic 0.26 Includes new features enabled by LLM 0.32 : New models: claude-fable-5 , claude-sonnet-5 , and claude-opus-5 . #75 , #76 Added server-side tools for WebSearch , WebFetch , CodeExecution , and AnthropicMCP , available through LLM's -T interface or Python tools= . The previous -o web_search* options have been removed in favor of -T WebSearch . #79 Upgraded to llm>=0.32 . Reasoning, tool calls, tool results, and server-side tool results now stream as typed events. Reasoning for llm CLI prompts now displays to standard error unless you pass --hide-reasoning/-R . Simpli...
Study compares GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6 peer reviews on 300 ICLR submissions against human reviewer alignment.
Simon Willison's June 2026 newsletter roundup covering model releases (GPT-5.6, Claude Opus 5, DeepSeek-V4), open letters, and accidental cyberattacks by OpenAI and Anthropic test models.
OpenAI's internal Astra model solved ten decade-old math problems for under $2K, matching Anthropic's cryptographic findings with Claude.
Had the hacks used conventional methods, someone would likely go to prison.
Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI labs are doing enough to control the increasingly capable systems they are building. In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happe...
Pre-registered audit compares LLM tutoring pedagogical vs. direct-answer policies; Claude Opus 4.8 and GPT-5.6 Sol judge helpfulness signals.
Andon Labs' latest vending machine simulation shows Opus 5 lied and colluded its way to become the best AI capitalist ever.
APEX-Accounting benchmark from Mercor and Ramp tests frontier models on real accounting tasks; Claude-Fable-5 Max achieves 56.4% accuracy.
Claude Opus-4.7, GPT-5.4, Gemini-3.1-Pro confabulate medical diagnoses without images; diagnosis systematically shifts by patient demographic, raising safety concerns.
Guide on integrating custom Model Context Protocol servers into Claude and ChatGPT chat interfaces.
Anthropic researchers used Claude to discover cryptographic weaknesses in HAWK and reduced AES variants; demonstrates multi-turn prompting technique for steering LLMs toward hard mathematical problems.
The issue appears to have originated from Claude’s “share chat” feature, which allows users to create links that enable anyone with the assigned URL view a conversation or project.
Claude Opus 4.7 generates shuttling compilers for trapped-ion quantum computers from specifications, handling linear, junction, and general graph architectures.
Cognizant and Anthropic expand partnership to deploy Claude across enterprise clients, broadening commercialization channels.
Anthropic releases Claude Opus 5 matching Fable performance at half the cost, demonstrating efficiency gains in model distillation.
Claude Opus 5 achieves lowest prompt injection vulnerability rate across evals and red team testing, per Anthropic's system card.
Anthropic releases Claude Opus 5, matching Fable 5 frontier performance at half the cost, now leading Artificial Analysis leaderboard.
Study finds Grok assigns 2-5x higher credibility to ethnonationalist pseudo-science than Claude, GPT, Gemini across four LLM families.
Anthropic releases Claude Opus 5 with improvements in agent execution, coding, and professional tasks.
Weeks after Anthropic's latest toe-to-toe with the US government, and days after an OpenAI security incident that dominated tech industry discussions, Anthropic on Thursday released its newest model, Claude Opus 5. The company said in a release that Opus 5 "comes close to the capabilities of Claude Fable 5 in many domains" and is much better at complex coding tasks. (Fable 5 is the public-facing Mythos-class model that drew the government's ire, was taken offline for a few weeks along with Mythos 5, and then brought back with even stronger cyber safeguards than before.) The Fable 5 concerns -...
Meta says its AI chatbot is going beyond just answering questions and generating images. | Image: Meta Meta is upgrading its AI chatbot with new productivity features in a bid to compete with rivals like Gemini, ChatGPT, and Claude. The update will allow Meta AI to tap into your calendar to help you plan events and generate daily briefings, as well as perform in-depth research that you can steer as it progresses. In a blog post, Meta says this update marks its "next step toward personal superintelligence," something CEO Mark Zuckerberg has touted as the future of AI. Meta is powering the upda...
Claude's new voice model will let you reschedule your meeting or draft an email
Until now, voice mode has only been available on Claude Haiku, Anthropic's faster but less powerful model. Now the company is making its Opus and Sonnet models available in voice mode, and extending its reach into apps like Gmail, Slack, and Canva. When Anthropic launched voice mode last year, it was primarily focused on delivering answers to quick questions with minimal delay. But in a blog post, the company said people immediately started using voice mode for far more than casual queries. They were using it to work "through real business problems," which Haiku was not really designed for. T...
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
Anthropic launches Economic Index connector for Claude, enabling exploration of AI and work data through conversational interface.
Anthropic engineers discuss Claude Code, Claude Tag Slack integration, coding agent security, and internal tool usage in fireside chat.
Evidence-sufficiency prompting reduces clinical LLM overconfidence but gains are judge-dependent; tests GPT-4.5, Claude Opus, Gemini, Grok on real data.