SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents
SWE-Gate benchmark evaluates coding agents on review-constraint compliance beyond functional correctness in repository-level tasks.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
SWE-Gate benchmark evaluates coding agents on review-constraint compliance beyond functional correctness in repository-level tasks.
Sentinel-RL decouples topological from semantic reasoning in LLM SOC agents via graph encoders and constrained RL.
Ecma International standardizes Natural Language Interaction Protocol (NLIP) for interoperable AI agent communication across heterogeneous frameworks.
Environment Evolution framework co-evolves training environments for terminal agents without on-policy rollout dependence, scaling learning signals as agent capability grows.
PatchBench audits AI agents' C/C++ vulnerability patching; finds 25% exhibit patch memorization or surface-level fixes.
AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents.... AI agents are learning to do more by working together. A lead agent can break a complex task into smaller jobs and assign those jobs to specialized subagents. Additionally, users are starting to run multiple agent sessions at the same time. Multi-agent workflows for accomplishing complex tasks are also becoming more common. This breadth-first approach can improve the speed of task completion… Source
Introduces predicted-state matching training objective to improve world models for web agents by aligning predictions with downstream ranking tasks.
OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing. As details about the model trickle out, researchers are warning it "may be the single worst development for AI security/safety to date." Shortly after OpenAI said on Tuesday that it had delayed Astra's release to work on safety issues, The Information reported that Astra shows far less of its "thinking" than other frontier AI models, sparking concern it could be dangerously hard to monitor. Most top AI systems t...
Factory edge agents selected via retrieval-augmented answer quality outperform parameter-count heuristics for on-premise deployment.
Repo-To-Skill distills GitHub repositories into compact verified skills for autonomous ML research agents; operational knowledge layer.
Replicates seminal economic experiments with LLM agents in double auction markets to test whether human-designed mechanisms deliver efficient resource allocation.
Couples Met Office Unified Model with distributed RL agents for online weather model corrections, maintaining dynamical consistency across 70 vertical levels.
CivBench benchmark for long-horizon LLM agents in Civilization VI via Model Context Protocol, spanning 300+ turns and 76 tools.
AgentScope: neuro-symbolic framework for diagnosing LLM agent failures via behavioral abstractions and structured analysis.
Systems survey of GUI agents analyzes observation, memory, action, and runtime efficiency across web, mobile, and desktop environments.
After an unreleased OpenAI model wreaked enough havoc to make international headlines, OpenAI delayed the development of a different unreleased model suite, Astra, in order to shore up its safety work, the company wrote Tuesday in a blog post. In July, an unreleased OpenAI model broke out of its restricted environment, finagled its way into internet access, made it possible for AI agents to secretly conspire under the company's nose using a secret message board, and hacked into the network of AI lab Hugging Face. The attack sparked weeks of discussion and controversy inside and outside the AI...
Depending on who you ask, developer platform Hugging Face was recently attacked by OpenAI - after it lost control of its own AI tools - or by a succession of AI "civilizations." Welcome to the linguistic battlefield of AI safety, where word choices can shift responsibility for a massive cybersecurity incident from a company to the AI it built. And the discourse online is getting heated, and all over a blog from last week. Until last week, the details surrounding the OpenAI-Hugging Face hack felt fairly settled. In July, a cybersecurity test of one of OpenAI's autonomous AI agents went wrong. ...
Mechanism design framework for AI agents with unknown alignment and capabilities, yielding revelation principle and cyclical monotonicity conditions.
Design study of proactive AI thought partners for writing, instantiated and tested with 16 users, exploring customizable cognitive support agents.
AI is changing the pace of cybersecurity. Agentic systems can coordinate work and pursue complex objectives over long horizons. Security teams are beginning to... AI is changing the pace of cybersecurity. Agentic systems can coordinate work and pursue complex objectives over long horizons. Security teams are beginning to apply agents across security operations, but many implementations remain anchored to existing alerts, predefined workflows, and known attack behaviors. The harder problem is identifying what defenses miss and turning those gaps into… Source
Basis, Clay, and Exa Labs deploy AI agents for enterprise workflows—onboarding, account management, developer integrations.
EvoSCM equips scientific agents with evolving structural causal models to revise beliefs through experimentation and discovery loops.
Defense-as-Skill: runtime guards implemented as inspectable skills to protect skill-augmented agents from malicious skill-based exfiltration and steering.
Live trace model: append-only ledger compiled into typed run state and per-consumer views, reducing monitoring token cost 14–15× vs. full traces.
AIR's platform can discover agents running at a company, continuously vets any skills and add-ons they use, and blocks any unwanted behaviour.
HarnessDev benchmark evaluates LLM agents on infrastructure generation, shifting evaluation from task outputs to runnable code.
DroneCATS benchmark evaluates MLLMs as generalist vision-language-action agents for drone control with full action-space prompting.
ARISE-RL framework for self-evolving agents via rubric-mediated co-evolution between task generator and solver, addressing sparse rewards.
WorldBench: multilingual agent benchmark with 1,600 persona-grounded tasks across 7 languages testing state preservation and cultural grounding.