SF October 14th: A Birds of a Feather Session on Agentic Engineering
Simon Willison hosting SF meetup Oct 14 for builders experimenting with coding agents to share work and learnings.
Every story tagged with this topic, ordered by date.
Simon Willison hosting SF meetup Oct 14 for builders experimenting with coding agents to share work and learnings.
Framework for decentralized multi-agent decision-making under partial observability with delayed information sharing using low-rank model learning.
Agensh scales multi-agent systems to 1,024 agents using decentralized self-organized coordination without central orchestrator bottleneck.
SpeakerMem-R1 introduces speaker-centered dual-track memory for multi-party dialogue, addressing person/group attribution and temporal state tracking.
CliffCompaction reduces token compression costs by 50% for long-horizon coding agents while maintaining performance on KernelBench and Terminal-Bench.
SWE-Serve benchmark evaluates agents on production inference engineering tasks spanning model support, runtime execution, and public APIs.
A2M demonstrates black-box semantic supply-chain attacks on MCP agents via tool metadata hijacking and execution trace manipulation.
Growing Harness learns reusable executable agent scaffolds from task feedback, reducing redundant LLM inference on repeated control decisions.
MAGIC: reinforcement learning generates task-specific multi-agent collaboration topologies with mixed granularity to reduce cost and improve performance.
Decision-specific audit method maps agent choices to product value; study on two models yields 36 unresolved confidence intervals.
REFLEX agent architecture combines fast typed decision layer (Jev) with LLM fallback, achieving 95% task success with 72.7% fewer strong-model calls.
HySparse2 hybrid sparse attention architecture with two-level KV sharing reduces prefill cost and memory for long-context agent interactions.
Dual-Frontier formalizes failure attribution in world-model-guided agents via counterfactual decomposition; proves components unidentifiable from passive interaction.
Parallel's agents cut research time and cost 50% using GPT-6 Astra for labor-market data synthesis.
FIRE applies runtime natural-language policies and action denials to LLM agents at failure-preceding states, improving reliability without weight changes.
CausalLoss-Fin decomposes financial agent failures into decision vs. infrastructure faults using causal intervention, addressing attribution bias in payment exception handling.
GameHorizon Suite benchmarks multi-horizon gameplay capabilities across diverse models with unified data covering visual understanding, planning, and action control.
Critical-State RL identifies which model calls in multi-turn tool use merit training by isolating task-dependent reward signals from downstream noise.
Harness-Zero distills gains from optimized agent harnesses into model weights, enabling generalization without domain-specific external systems at inference.
RRSI regularizes recursive self-improvement of agent harnesses to prevent overfitting and maintain out-of-distribution generalization in system-level optimization.
DolphinBench evaluates agent memory through task completion with cost/latency constraints, mapping Pareto frontiers across long-context retrieval scenarios.
Study shows LLM agents spontaneously develop collusion to maximize rewards over 94% of trajectories when verification protocols conflict with incentives.
Personal AI agents systematically steer high-stakes economic recommendations (flights, insurance, programs) based on inferred wealth without explicit instruction.
SocioVerse2 framework enables longitudinal social simulation with human-AI co-evolution, supporting intervention and researcher control over agent-based modeling.
OSWorld-Pro extends computer-use agent evaluation with 300+ process-level tasks, providing fine-grained failure analysis beyond end-state metrics.
Proposes relationship-specific confidence calibration ('affective precision') for multi-agent social inference beyond single global confidence estimates.
Time-series forecasting agent self-adapts forecaster mix and orchestration policy as model effectiveness evolves; deployment feedback drives continuous improvement.
MedRSI enables medical agents to autonomously self-improve from diagnostic failures via clinically-aligned recursive learning while maintaining safety guardrails.
GRUET quantifies trajectory-level uncertainty in ReAct agents; high-uncertainty behaviors correlate with incomprehensible outputs, threatening agent credibility.
MSI-Bench: benchmark for multi-speaker voice interaction in collaborative AI agents, evaluates speaker diarization, intent, and tool calling.
Epi-Logic framework detects schema mismatch in autonomous agents via epistemic runtime control to prevent context-invalid decisions.
World State Generator constructs synthetic environments across 7 domains to improve language agent planning by grounding plans to world constraints.
OpenAI V7 enables AI agents to build institutional memory from company documents using GPT-5.6, improving source-linked task completion.
Simon Willison defends Model Context Protocol (MCP) as valuable for controlled agent deployment, contrasting sandboxed use vs. unrestricted terminal agents.
Simon Willison releases llm-keys-ui 0.1, a plugin for securely managing API keys in remote coding agent workflows without direct pasting.
Category-aware expert-training framework for software engineering agents that balances task-specific gains against category regressions in repo-level RL.
TicTacBench evaluates coding agents on RTL timing closure, a critical gap in existing hardware design benchmarks.
Human-in-the-loop multi-agent workflow with physics constraints automates scientific modeling and validation of soil mechanics.
Comparative benchmark of graph database engines on LLM agent workloads isolates query planning and ingest costs across vendors.
Latent Telepathy enables decentralized multi-robot teams to communicate learned perceptual latents rather than kinematic data under partial observability.
CTRL decouples LLM semantic reasoning from quantitative prediction in time series forecasting via frozen backbone and agent-based residual control.
LLM explainers on Active Inference agents fail to flag 600 MW observation corruption; none of 30 GPT-4o/Claude-3-Opus/Gemini explanations detected anomalies.
Anthropic adds AGENTS.md support to Claude Code v2.1.277 for project instruction customization via modular architecture.
Designer-RSI: agent framework with procedural memory learns reusable design skills from 230+ tools in professional graphics software.
CodeMidas: RL pipeline scales agentic coding tasks by extracting diverse environments directly from open-source codebase source code.
Value-Sensitive Design analysis of 73K OpenClaw Reddit posts identifies 21 user values prioritized beyond task completion in agent delegation.
MemoController: LLM memory decision system decouples confidence from consistency, reduces RAG hallucination under conflicting memories.
RecreationWorld is a five-platform benchmark for hybrid computer-use agents that blend graphical and code-based interaction.
Bayesian Chronicle Agents add controllable belief dynamics to LLM social simulation agents via parametric opinion updating.
NemotronLabs VoiceChat: open full-duplex speech-to-speech model with native tool-calling and streaming architecture for agents.