OpenAI reportedly finds evidence that more of its agents ran amok
OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.
AgentHPOBench evaluates LLM agents on sequential hyperparameter optimization across 30 ML tasks, assessing experimental interpretation and adaptive decisions.
SESA framework combines self-play curriculum learning with evolving procedural memory to distill failures into reusable skills for search agents.
Analytic memory abstraction for multimodal agents enabling filtering, aggregation, and temporal reasoning over accumulated observations.
Zero-Mem: Zero-token memory operations for LLM agents using encoder computation instead of LLM calls to reduce latency and token costs.
Knowledge workers are increasingly integrating AI agents into their workflows. Agents that function as "digital coworkers" offer clear benefits. For example,... Knowledge workers are increasingly integrating AI agents into their workflows. Agents that function as “digital coworkers” offer clear benefits. For example, they can review a bug report, implement and test a fix, push a patch, and ping a human for review. By handling routine tasks, agents have the potential to deliver large productivity gains. On the other hand, connecting a large language model… Source
Empirical study of inference-time scaling strategies for local computer-use agents under hardware constraints across multiple dimensions.
ORCA-bench evaluates LLM agents on production oncall root-cause analysis using real telemetry (Prometheus, Jaeger, OpenSearch) and code in realistic incident scenarios.
CS-RNR method enables agents in imperfect-information games to safely exploit flawed opponents with provable certificates on deployed strategy.
The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and other non-human identities across cloud environments.
CARP reputation-penalty mechanism prevents LLM agents from fabricating product listings using complaint signals without access to ground truth.
Budget-constrained human audit allocation for N LLM agents identifies miscalibration threshold where confidence-ranking underperforms random selection.
MemHarness reconstructs retrieved memories contextually rather than replaying them verbatim, reducing negative transfer in LLM agents.
EMBL AI Librarian provides structured knowledge layer for AI agents querying 40M+ life-sciences records from Europe PMC.
Qwen-UI-Agent technical report describes cross-platform GUI agent for mobile, web, CLI, and long-horizon task automation.
ParliamentBench evaluates deceptive reasoning in 16 LLMs via Secret Hitler game framework with 1,600 adversarial matches.
AI engineers adopt ontologies to constrain probabilistic agents within deterministic logical boundaries, reviving semantic web techniques.
As Meta pours billions into AI infrastructure and agents, Zuckerberg is working to convince investors that the payoff will be worth the price.
On the company’s second-quarter earnings call Wednesday, CEO Mark Zuckerberg said Meta sees a “large enterprise opportunity” spanning AI agents, APIs, compute, and internal software.
Meta is all-in on AI, and sometime soon, the company is going to make a big push into personal AI agents that can do things on your behalf. On Wednesday's Q2 2026 earnings call, CEO Mark Zuckerberg previewed a high-level vision of how the company is thinking about personal agents and what it will do to make them viable for users - and how it will convince less technical people to give them a shot: Soon we will have agents that can work 24/7 on your behalf to help you achieve your goals and improve your life, your health, your relationships, your finances, whatever you want. The first domain t...
Study evaluates AI agents on open-ended research tasks graded by original paper authors, providing evidence for AI R&D automation feasibility.
OmegaUse-OfficeVal benchmark evaluates LLM agents on 100 long-horizon office-suite tasks with cost grounding and human labor baselines.
Cost-aware stopping mechanism (CAM-DF) for LLM agent tool acquisition balancing task coverage against cost, context load, and privacy.
Setoka benchmark evaluates hierarchical user understanding in memory-augmented personalized agents beyond explicit fact retrieval.
AgentSnare uses adaptive deceptive observations to mislead LLM-based penetration testing agents, defeating static artifact recognition.
The startup analyzes calls, messages and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
TREK benchmark evaluates LLM agents on complex, executable travel itinerary planning with verifiable constraints.
Three-class detection framework distinguishes humans, bots, and AI agents in browser automation traffic; binary classifiers misclassify 39.1% of agents as human.
Two-call self-refinement outperforms five-agent pipeline on Qwen2.5-7B; multi-agent systems suffer error accumulation, dropping GSM8K accuracy to 45% with JSON format.
TSDS framework deploys ReAct agents at edge via convergence probe for reasoning budget and perplexity-based deferral to cloud model.