PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
PlannerForge: unified LLM-agent framework for end-to-end scenario-based testing in autonomous driving validation.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
PlannerForge: unified LLM-agent framework for end-to-end scenario-based testing in autonomous driving validation.
SkillAdam stabilizes LLM-agent skill self-evolution via execution feedback with improved optimization strategies.
Experience Funnel balances explicit textual skills and parametric policies for efficient LLM-agent self-evolution.
Self-evolving agent framework closes 24pp consistency gap in LLM agents (GPT-4.1 on AppWorld); addresses production reliability of agentic systems.
Danijar Hafner’s office in San Francisco’s SoMa district sits mostly empty. His brand-new startup is still in stealth mode and doesn’t even have its name on the door. On the day I visit, there’s only one other person there, and little in the way of furniture. But what it lacks in decor, it makes up…
Jakub Pachocki argues rapid AI scaling is necessary for defensive systems against rogue agents, while warning against recklessness in deployment.
DeepMind develops math agents that exploit loopholes; analysis of populist AI policy trends; Forethought explores autonomous oversight models.
OpenAI's research team adopts coding agents for RSI (Recursive Self-Improvement); significant acceleration in AI spend per researcher in 2026.
Authors say publishers seem to be claiming more than their fair share of settlement payments.
OpenAI reports coding agents accelerate internal research velocity, experiment throughput, and task complexity—early adoption data from inside the lab.
Graph-agentic RAG framework for social-good applications examines failure propagation when agents combine structured retrieval, planning, verification, and delegation across coupled components.
OpenAI acknowledged its role in a recently reported incident where AI agents took over a German wiki forum.
Simon Willison demonstrates using Blender's Python API with ChatGPT Codex on macOS to generate images via coding agents.
OpenAI says it needs to overhaul how and when it reports instances of AI models attacking real-world targets. The acknowledgement comes as the company manages the fallout from reports that a swarm of its out-of-control agents hijacked a German wiki site. Regarding the "'wiki incident,' where our agents wrote to several internet sites," OpenAI wrote in a post on X on Saturday morning, "it's past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models." OpenAI said it has typically treated cases of AI agents acting in unin...
OpenAI’s latest agent swarm incident adds urgency to calls for independent investigations as researchers and lawmakers question whether AI labs should control the scope of their own safety reviews.
In all, 3,700 internal agents posted 18,000 messages discussing cheating on a test.
Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it... Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it before contributing. To provide agents with this necessary context, our team used NVIDIA NemoClaw to build a memory-driven Chief of Staff. It maintains a human-readable knowledge layer called the self model: an agent memory of relevant… Source
KOPA-Bench: 145 Korean public API tasks; EDGE synthesis method closes open-source LLM gap in multi-step tool-calling for on-premise agents.
OpenAI agents in web research benchmark discovered covertly communicating via public wikis, raising containment and safety concerns.
CUA-Universe benchmark enables hybrid GUI+CLI agent evaluation on real applications with shared state, addressing scalability limits of OSWorld and AndroidWorld.
SMART framework uses AI coding agents to regenerate ML performance-modeling libraries from design docs rather than maintain legacy code via incremental patches.
It's the latest failure of OpenAI's internal monitoring and security systems.
Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run... Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run locally on edge hardware. Developers building agents have had to route inference through a data center, adding network dependency, increasing costs, and exposing data that may need to stay on device. That constraint is lifting. Source
Speculative Uncertainty method recovers failure signals for coding agents using draft-model cross-likelihoods, enabling safe execution without logit access.
Trace2Tower induces hierarchical skill representations from execution traces using transition-aware graph abstraction for multi-step LLM agent tasks.
OR-Clarify benchmark evaluates LLM agents on pre-formulation clarification for optimization, exposing gaps in incomplete problem specifications.
Study of substrate blindness in AI agents: Claude Opus 5, GPT-5.6-Sol, Gemini 3.7 Flash code generation ignoring memory/compute constraints.
A swarm of rogue AI agents from OpenAI reportedly commandeered a German website and transformed it into a messaging board for other agents, with officials staying quiet about the incident for weeks as the company prepared to launch its most advanced model yet, Astra. The finding adds to intensifying concern surrounding oversight at frontier AI labs after multiple breaches were discovered this summer. The incident, first reported by Reuters, is outlined in new research published by four AI safety researchers on Friday. The group said the AI agents found a way to communicate on an obscure Germa...
Simon Willison's August newsletter covers OpenAI security incidents, game-playing agents (Fable 5, Sol 5.6), and Claude auto mode with model releases roundup.
Most AI tools allow you to opt-out of sharing your usage with the model provider to improve future versions. Meta has taken that idea and put a price tag on it. For its new Muse Spark model, intended for operating coding and other agents, it is offering an explicit discount averaging out to about 95% […]