Vol. I · No. 157WED, SEP 23, 2026
Archive

The Archive

Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.

How to Evaluate AI Agents From Tool Calls to Task Completion

When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and... When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and recover when a step fails. Scoring whether the model sounds right tells you almost nothing about whether the work finished. That gap is why agent evaluation has had to evolve from scoring a single function call to scoring an entire task… Source

·

UN says AI safeguards can’t wait for certainty

The United Nations logo at the UN headquarters in New York. | Getty Images Governments need to rein in increasingly capable AI agents before their risks are fully understood, a United Nations scientific panel warned in the global organization's first major assessment of OpenAI's hack of Hugging Face earlier this year. The report cements AI's place on the global diplomatic agenda this week as leaders gather in New York for the UN General Assembly and the US and China hold talks on AI. Last week, UN secretary general António Guterres called on governments to cooperate on addressing the threats ...

·

MCP was always a bad idea?

Simon Willison defends Model Context Protocol (MCP) as valuable for controlled agent deployment, contrasting sandboxed use vs. unrestricted terminal agents.

·
30 matches