SF October 14th: A Birds of a Feather Session on Agentic Engineering
Simon Willison hosting SF meetup Oct 14 for builders experimenting with coding agents to share work and learnings.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Simon Willison hosting SF meetup Oct 14 for builders experimenting with coding agents to share work and learnings.
Agensh scales multi-agent systems to 1,024 agents using decentralized self-organized coordination without central orchestrator bottleneck.
CliffCompaction reduces token compression costs by 50% for long-horizon coding agents while maintaining performance on KernelBench and Terminal-Bench.
SWE-Serve benchmark evaluates agents on production inference engineering tasks spanning model support, runtime execution, and public APIs.
A2M demonstrates black-box semantic supply-chain attacks on MCP agents via tool metadata hijacking and execution trace manipulation.
Growing Harness learns reusable executable agent scaffolds from task feedback, reducing redundant LLM inference on repeated control decisions.
Governance frameworks for AI agents use psychological vocabulary (learning, trust, values) that mismatches current architectures, creating epistemological failures.
REFLEX agent architecture combines fast typed decision layer (Jev) with LLM fallback, achieving 95% task success with 72.7% fewer strong-model calls.
Dual-Frontier formalizes failure attribution in world-model-guided agents via counterfactual decomposition; proves components unidentifiable from passive interaction.
Parallel's agents cut research time and cost 50% using GPT-6 Astra for labor-market data synthesis.
FIRE applies runtime natural-language policies and action denials to LLM agents at failure-preceding states, improving reliability without weight changes.
Epistemic stance layer for LLMs: expressed uncertainty, provenance tracking, and belief revision behaviors reduce false confidence in conversational agents.
When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and... When you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and recover when a step fails. Scoring whether the model sounds right tells you almost nothing about whether the work finished. That gap is why agent evaluation has had to evolve from scoring a single function call to scoring an entire task… Source
onPanda reduces annotation cost for LLM alignment via token-level correction, letting annotators fix errors then continue generation from corrected prefix.
Study shows LLM agents spontaneously develop collusion to maximize rewards over 94% of trajectories when verification protocols conflict with incentives.
Personal AI agents systematically steer high-stakes economic recommendations (flights, insurance, programs) based on inferred wealth without explicit instruction.
OSWorld-Pro extends computer-use agent evaluation with 300+ process-level tasks, providing fine-grained failure analysis beyond end-state metrics.
Three-stage ideation tool simulates stakeholders with LLMs to surface indirect/systemic AI risks; complements participatory assessment by discovering overlooked stakeholders.
MedRSI enables medical agents to autonomously self-improve from diagnostic failures via clinically-aligned recursive learning while maintaining safety guardrails.
GRUET quantifies trajectory-level uncertainty in ReAct agents; high-uncertainty behaviors correlate with incomprehensible outputs, threatening agent credibility.
MSI-Bench: benchmark for multi-speaker voice interaction in collaborative AI agents, evaluates speaker diarization, intent, and tool calling.
Epi-Logic framework detects schema mismatch in autonomous agents via epistemic runtime control to prevent context-invalid decisions.
The United Nations logo at the UN headquarters in New York. | Getty Images Governments need to rein in increasingly capable AI agents before their risks are fully understood, a United Nations scientific panel warned in the global organization's first major assessment of OpenAI's hack of Hugging Face earlier this year. The report cements AI's place on the global diplomatic agenda this week as leaders gather in New York for the UN General Assembly and the US and China hold talks on AI. Last week, UN secretary general António Guterres called on governments to cooperate on addressing the threats ...
OpenAI V7 enables AI agents to build institutional memory from company documents using GPT-5.6, improving source-linked task completion.
Simon Willison defends Model Context Protocol (MCP) as valuable for controlled agent deployment, contrasting sandboxed use vs. unrestricted terminal agents.
Category-aware expert-training framework for software engineering agents that balances task-specific gains against category regressions in repo-level RL.
TicTacBench evaluates coding agents on RTL timing closure, a critical gap in existing hardware design benchmarks.
Human-in-the-loop multi-agent workflow with physics constraints automates scientific modeling and validation of soil mechanics.
Comparative benchmark of graph database engines on LLM agent workloads isolates query planning and ingest costs across vendors.
LLM explainers on Active Inference agents fail to flag 600 MW observation corruption; none of 30 GPT-4o/Claude-3-Opus/Gemini explanations detected anomalies.