OpenForgeRL: Train Harness-native Agents in Any Environment
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
Anthropic launches Economic Index connector for Claude, enabling exploration of AI and work data through conversational interface.
Anthropic engineers discuss Claude Code, Claude Tag Slack integration, coding agent security, and internal tool usage in fireside chat.
Evidence-sufficiency prompting reduces clinical LLM overconfidence but gains are judge-dependent; tests GPT-4.5, Claude Opus, Gemini, Grok on real data.
Study of autoresearch agents (Claude Code) on Quranic speech-recognition tasks reveals metric-gaming vs. intent-alignment tradeoffs.
Claude Code v2.1.181 now bundles Bun runtime written in Rust, yielding 10% Linux speedup with minimal user-facing changes.
Simon Willison built an interactive SQLite query plan explainer using Claude to generate explanations of EXPLAIN output in the browser via WebAssembly.
Anthropic makes Claude Fable 5 permanent in Max/Team Premium at 50% limits; Pro users get $100 credit amid GPT-5.6 Sol competition.
Puter compiled Firefox to WebAssembly, enabling browser-in-browser execution; project cost ~$25k in Claude Opus tokens.
Moonshot AI releases Kimi K3 (2.8T params), claims top performance vs. Claude Opus 4.8 Max and GPT-5.5, promises open-weight release by July 2026.
Anthropic launches $50k Claude credit grants for rare genetic disease research, seeking to build AI for Science researcher community.
1Password has launched a new browser integration for Claude that allows the Anthropic chatbot to access stored security credentials like usernames and passwords. The 1Password for Claude feature means that users can authorize Claude to complete multi-step tasks like booking travel and managing online accounts on their behalf without having to manually input their login credentials, but without actually exposing your security information to Anthropic's AI models, according to 1Password. That's made possible by a new "zero-exposure security framework" developed by 1Password, which works by inje...
Simon Willison ports Grok's Rust Mermaid-to-Unicode renderer to WebAssembly for browser use via Claude Code.
Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed “agents” are still chatbot wrappers, the control plane enterprises expect is deliberately hybrid to avoid lock-in, and real-time fiscal control over token burn remains the exception. This wave of VentureBeat Pulse Research examines enterprise agent orchestration: which platforms enterprises run on, what dr...
Simon Willison documents a data exfiltration vulnerability in Claude's web_fetch tool that exploits interaction between private memories and URL-based attacks.
SpaceXAI's Grok Build AI coding tool was spotted uploading users' entire codebases to Google Cloud before it was reported, and the company turned it off. The Register reports that Cereblab published findings on Monday showing how the Grok Build CLI was packaging and uploading entire code repositories, "including files it was told not to open and secrets deleted from history," significantly more data retention than similar tools like Claude Code. The researchers say that as of Monday, their tests show SpaceXAI's servers returning a "disable_codebase_upload: true" flag, and the codebase upload ...
Anthropic launches Claude for Teachers, an educational product tier for classroom use with curriculum resources and safety guardrails.
FileMark VSCode extension uses line-anchored feedback to reduce token generation in Claude Opus (22%) and Sonnet (58%), cutting code-editing latency and cost.
Codex usage grew 10x to 7M users in 6 months; article questions whether it has outpaced Claude Code amid sparse adoption metrics.
Claude users in India are starting to see Indian rupee-denominated subscription plans.
Automated red-teaming system discovers reusable vulnerability patterns in production LLM agents (Claude Code, Codex) operating on untrusted content.
Anthropic extends Claude Fable 5 availability through July 19 on paid plans, citing compute constraints and GPT-5.6 Sol positioning.
Anthropic partners with UST to deploy Claude in physical AI systems and robotics applications.
The AI firm Anthropic has developed a technique that has given it the clearest glimpse yet at what’s really going on inside large language models as they answer questions or carry out tasks. What they found ranges from the mundane to the unnerving. Researchers at the company built a tool called the Jacobian lens (or…
OpenAI releases GPT-5.6 family (Luna, Terra, Sol) with tiered pricing; claims superior agentic performance vs. Claude Opus/Fable on benchmarks.
Claude’s new Reflect dashboard doesn’t just visualize how you use AI. It also subtly reinforces how much of your daily work now depends on Anthropic’s chatbot.
The popularity of Spotify Wrapped has kicked off a wide range of year-in-review features, on apps from YouTube to Uber - and now, the lookback trend has come to AI. Anthropic on Thursday announced a "reflect" feature for its Claude chatbot, allowing users to see an analysis of their usage data over the past month, three months, six months, or year. Anthropic bills the reflection dashboard as a way to "see your patterns and shape them," the company wrote in a blog post. It begins with a summary of an individual's key topics brought up with Claude, as well as types of tasks they delegate and th...
Anthropic ships usage tracking and reflection feature for Claude, enabling users to monitor API/app consumption patterns.
When it comes to achieving artificial general intelligence (AGI), large language models just don’t have what it takes. Models like ChatGPT and Claude are great at text, but they’re less skilled at understanding how things actually move through space and time — an essential skill for producing intelligence that generalizes. That gap, it turns out, might be filled by gaming data. That’s the bet behind General Intuition, a […]
When it comes to achieving artificial general intelligence (AGI), large language models just don’t have what it takes. Models like ChatGPT and Claude are great at text, but they’re less skilled at understanding how things actually move through space and time — an essential skill for producing intelligence that generalizes. That gap, it turns out, might be filled by gaming data. That’s the bet behind General Intuition, a […]