Gemini API Managed Agents: 3.6 Flash, hooks, and more
Google ships Gemini API Managed Agents with Gemini 3.6 Flash and tool-use hooks for production agent deployment.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Google ships Gemini API Managed Agents with Gemini 3.6 Flash and tool-use hooks for production agent deployment.
Messier consolidates 957K evaluation records across 30 benchmarks and 714 agents to enable comparable cross-benchmark agent assessment.
Custom harness engineering distributes security controls for AI coding agents without vendor lock-in, enabling organizational scaling.
RSIBench-Data isolates LLM agents' data-centric research capability in post-training loops, decoupling research from systems optimization.
OpenAI's Akshay Nathan details ChatGPT Work product strategy: Sites, memory, subagents, finance, no-code tools scaling from 0 to 10M users.
HiSkill hierarchical skill graph framework organizes LLM agent trajectories into structured graphs linking high-level skills to atomic operations for long-horizon tasks.
Joint agent-speculator RL aligns next tool-call prediction with deployed agent behavior by unifying speculator and agent in single model, reducing latency.
WorkSurface-Bench evaluates enterprise agents on multi-source knowledge routing across documents, tables, and graphs with 1,151 tasks.
HYSET evaluates tool sets jointly for LLM agents via hyperedge prediction, improving on sequential tool retrieval.
Perplexity has expanded its agentic Personal Computer tool to Windows, allowing computers running the world's most popular OS to be used as a locally run AI system. Like the Mac version that Perplexity launched in April, Personal Computer for Windows operates like a "general-purpose digital worker" that can access local files and apps to perform actions on your behalf, such as creating documents and updating spreadsheets. This launch builds on Personal Computer integrations that Perplexity launched for Microsoft's 365 workspace apps and Teams virtual meeting software in May. Personal Computer...
Multi-turn long-horizon planning for foundation agents via controlled pre/post-training with single/multi-teacher on-policy distillation.
APPA: IFC framework for LLM agents handling mixed-confidentiality data via context branching and prospective access control against injection attacks.
Study of correctness degradation in code repair agents under forced revision loops; 30 HumanEval benchmarks show revision doesn't guarantee reliability.
SIREN applies LLM agents to end-to-end extreme-weather early warning with multi-step decision workflows.
Evaluates fuzz testing methods for safety-critical RL agents across robotics and autonomous systems with standardized metrics.
Robotics learns from bitter lesson scaling; AI agents complete week-long programming tasks; OpenAI discovers accidental vulnerability in reasoning model.
Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in its domain. But they all have their own distinct knowledge and objectives. Today they can exchange data, but they are not yet able to actually coordinate…
For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflows, data, and systems. The platform best-suited to run agents is built with proper CPU capacity, resilient data access, policy-aware tool use, observability, memory management, and the…
Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware... Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware knowledge, precise reasoning, and repeated interaction with electronic design automation (EDA) tools. LLMs have accelerated code generation, and AI agents extend their impact by using verification feedback to iteratively correct errors. Source
Case study: Cohere's North AI agents reduce friction for wealth managers through task automation and time savings.
Skill Self-Play reconciles task diversity and verification reliability in LLM self-evolution via co-evolving skill agents.
Regression Tax study quantifies performance degradation when adding skills to LLM agents across 6K runs on office automation tasks.
Dynamic capability scoping framework for enterprise AI agents using three-source permission architecture to enforce least-privilege credential access and reduce attack surface.
Self-calibrating framework for LLM agents to detect and correct operational drift in open-ended autonomous edge resource allocation.
SceneActBench: benchmark for evaluating VLM agents' ability to act on complete multi-object 3D scenes via unified agent-environment loop.
ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents.
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
Open-source evaluation framework for open-weight LLM agents on longitudinal data tasks, addressing privacy constraints in research deployments.
VoLN: vision-only navigation benchmark and method for embodied agents without language instructions in GPS-denied environments.
Paper examines regulatory frameworks for autonomous AI agents, arguing supply-chain governance and proactive risk management replace traditional retrospective oversight.