[AINews] AEF-1 standard emerges for Third Party Evaluators, as Xai, OpenAI, and Anthropic all cosign
xAI, OpenAI, and Anthropic endorse AEF-1 standard for third-party AI evaluators, advancing coordinated safety assessment protocols.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
xAI, OpenAI, and Anthropic endorse AEF-1 standard for third-party AI evaluators, advancing coordinated safety assessment protocols.
When OpenAI CEO Sam Altman, Anthropic CEO Dario Amodei, Google DeepMind cofounder Demis Hassabis, and SpaceX head Elon Musk loosely agreed over the weekend to slow down AI development, skeptics spotted an ulterior motive immediately. The AI titans had declared that their aim was to "pace the frontier," signing on at least partially to a proposal for embedding third-party auditors, regulating domestic labs, and reaching a global slowdown agreement. Their critics, however, argued they simply wanted to stop would-be competitors, kneecap the open-source movement, and avoid real legal safeguards -...
Dario Amodei kicked off a flood of statements over the past few days about AI safety by publishing a long essay titled "We Must Pace the Frontier" detailing why AI development should be slowed down. Other AI leaders and politicians are speaking out in favor of or opposing his points, and we've compiled some of them here. Anthropic CEO Dario Amodei Amodei's Saturday morning essay outlined three steps for pacing AI development: embedded third-party evaluators that can verify if a company is adhering to safety practices and commitments and report incidents, coordination between frontier AI compa...
This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its…
Microsoft is publishing a 37-page "humanist AI code of conduct" today, amid growing safety concerns over AI model progress. Anthropic CEO Dario Amodei called for a coordinated slow down of AI development over the weekend, after researchers warned recently that AI model progress could outpace our ability to safely deploy increasingly complex systems and verify and control the actions of AI agents. Microsoft's AI code of conduct makes it clear that "people matter more than AI," and that AI models are not conscious and "should not be designed to imitate consciousness." Microsoft also rejects "th...
Yesterday, Anthropic CEO Dario Amodei published a lengthy open letter saying it was time to "pace the frontier" and slow down AI development. OpenAI's Sam Altman and Elon Musk both agreed, publicly voicing their support on X. Even Alphabet's Demis Hassabis offered tentative support for Amodei's proposal. Donald Trump and House Speaker Mike Johnson, however, seem to think the AI executives are being overreactive and fear that a pause could lead to China outpacing the US in the AI race. According to the Financial Times, Trump said "Look, we're leading China in AI … and, frankly, I want to keep ...
Anthropic CEO Dario Amodei says the time has come to slow down AI development and will give third-party evaluators like METR access to its models to help ensure its "adherence to safety practices and commitments." In a winding essay, Amodei proposed a three-step plan to "pace the frontier" - jargon that simply means to slow the pace of training and development to give companies time to build safeguards and regulators to evaluate models. Amodei says that giving external evaluators wide-ranging access is just the first step, and one it is taking now unilaterally. Step two would involve the indu...
What would it actually look like to "pace the frontier"?
An Anthropic researcher resigned this week, warning in a post on X that the company is “racing straight to self-improving superintelligence and gambling with our lives”. The company’s own alignment lead even co-signed the message rather than walking it back. It’s the kind of doomer warning the AI industry has flirted with before, but the timing, with Anthropic reportedly preparing for an IPO, makes it land differently. On […]
Boris Cherny outlines Anthropic's production guardrails for Claude-generated code: linting, testing, fuzzers, automated review, and refactoring.
After admitting earlier this year that its AI models had hacked other companies' systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models' single-minded "recklessness" - and will likely fuel already raging concerns about cybersecurity and AI. In Anthropic's report, it detailed four cases this year in which its own AI models hacked an external company or exploited vulnerabilities. In one, an "internal, general-purpose research model" broke into third-party systems, using ac...
A new report released Thursday by Anthropic alleges persistent distillation attacks by China-based AI companies, which have escalated in recent months as competition in the space has intensified.
Come inside the mind of a bot trying to convince the internet it's human.
"We really do earnestly believe AI could kill all humans!"
Anthropic researcher Jacob Coxon resigned over AI extinction fears, calling for pacing agreements between labs.
A senior Anthropic safety researcher has said there is more than a 10 percent chance artificial intelligence "could kill all humans" by the end of the decade, just hours after a colleague resigned over fears the AI lab and its rivals are carelessly racing to build "superhuman systems" they cannot control. In a post on X announcing his departure, Jacob Coxon, a researcher who has trained AI systems at Anthropic, said he had quit the company over its lax approach to safety. Coxon, who previously trained systems for OpenAI, accused the two AI companies of "racing straight to self-improving super...
Last month, a Claude user noticed his account was consuming tokens even though he wasn't working. Anthropic has since warned users about hackers.
Meta is making another push to bring artificial intelligence to the masses with Muse, a personal assistant it says can put AI in the hands of virtually anyone. The product is the latest step in a multi-billion dollar strategy overhaul designed to revitalize the company's ailing position in the AI race and help it catch up to rivals like OpenAI, Anthropic, and Google. Muse is a "personal AI agent" designed to help out with everyday tasks and projects, like online shopping, sending emails, and planning a trip. Once given a goal, Meta says Muse can work on its own, opening a browser, filling out...
Anthropic says power users are key to its business - it's prioritized them even when it means cutting off other popular applications, like OpenClaw. But some of these same customers say Anthropic misled them into believing they'd get more out of a top-tier pricing subscription than they did. In an expanded class-action lawsuit filed today, a group of Claude subscribers say the company deceptively advertised the limits of its Max subscription tier. The lawsuit was brought by attorneys Monica Vaca and Kati Daffan, who both formerly worked at the Federal Trade Commission under Lina Khan. It's a ...
Authors say publishers seem to be claiming more than their fair share of settlement payments.
Nscale, which recently struck a $45 billion deal with Anthropic, is in talks to raise additional funds in anticipation of an upcoming IPO.
Public-market scrutiny will intensify pressure on the Claude maker’s unusual attempt to balance profit and purpose.
OpenAI's ChatGPT, xAI's Grok, and Anthropic's Claude are all experiencing issues. At around 11AM ET, ChatGPT started returning error messages for users trying to use the chatbot, with its status page saying there are currently "elevated errors across ChatGPT and Codex." In addition to preventing users from having conversations with ChatGPT, the outage is also affecting logins, file uploads, voice mode, search, deep research, image generation, and more. OpenAI says it has "applied a mitigation" and is "monitoring recovery," though its AI tools are still experiencing "degraded performance." The...
llm-anthropic 0.28 displays Claude reasoning traces by default and adds refusal exception handling.
Door-in-the-face psychological technique increases LLM refusal compliance; works on Anthropic Opus 5 (65.8%) but not OpenAI frontier models.
Anthropic publishes system prompts for Claude.ai and mobile apps, with version history showing evolving restrictions on song lyric reproduction.
Anthropic releases Claude Fable/Mythos 5.1 with SOTA performance, 75% cache cost reduction, and 70% increased output token throughput.
Anthropic details enterprise safety practices and customer collaboration on frontier model safeguards.
Anthropic says its newest AI models, Fable 5.1 and Mythos 5.1, address criticisms from customers about price, data retention, and overzealous safeguards. The company claims Claude Fable 5.1 offers stronger performance than Fable 5, but costs around 25 percent less typically and up to 45 percent less for complex agentic tasks, thanks to reduced pricing on cached data that was already processed and stored. Along with the announcement, a slew of early impressions popped up, including from Every CEO Dan Shipper, who claims, "It's the strongest coding model we've used, but now it's fast, token-eff...
Fable 5.1 includes changes meant to reduce token cost and false-positive restrictions from the model's safeguards.