Autoresearch: The feedback loop behind self-improving agents
Introspection co-founder explains autoresearch loops, agent recipes, and self-improving systems while arguing humans remain essential to AI software development.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Introspection co-founder explains autoresearch loops, agent recipes, and self-improving systems while arguing humans remain essential to AI software development.
Cursor's Forward Deployed Engineers help enterprises implement AI agents as software factories, per Pauline Brunet.
Audit exposes reliability issues in coding-agent benchmarks (GSO, SWE-Perf, SWE-fficiency): runtime instability, scoring artifacts, selection bias.
Cloudflare is giving AI companies until September 15 to separate web crawlers used for search from those used for AI training and agents, or risk being blocked by default on many publisher sites.
Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer... Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer reinforcement learning with verifiable rewards (RLVR) workflows for reasoning and agent tasks. RL is now becoming a practical technique for specialized AI where enterprises need more accurate agents for domain-specific workflows. Source
DiscoPER framework enables open-ended autonomous scientific discovery via iterative meta-reflection and cross-finding synthesis in LLM agents.
Case study on software engineering with frontier AI coding agents reveals shift from implementation scarcity to governance, inspection, and maintainability challenges.
OpenAgent formalizes generalization gaps in LLM tool-use agents across query, action, observation, and domain shifts via controlled sandbox evaluation.
Test-time control framework DART-VLN mitigates memory decay and loop inefficiencies in vision-language navigation agents.
Study shows moderate LLM agent personality expression outperforms extremes on trust and goal-adoption in conversational behavior-change tasks.
SWE-Doctor: LLM-based code agent using multi-faceted bug reproduction tests for runtime diagnosis to improve software patch generation.
SEA architecture enables safe self-modifying agents by freezing base models and gating adaptations through anytime-valid certificates against error budgets.
Anthropic’s Claude Sonnet 5 brings stronger agentic capabilities, lower pricing, and improved safety, positioning the model as a cheaper alternative to Opus, GPT-5.5, and Gemini Pro.
Anthropic releases Claude Sonnet 5, a frontier model optimized for coding, agents, and professional workflows at scale.
QVal proposes efficient evaluation method for dense supervision signals in long-horizon LLM agents without expensive end-to-end training.
Generative Skill Composition uses LLM-guided retrieval to select and compose skills for complex agent tasks without exposing full skill library.
Startup Acti is betting the smartphone keyboard is the next home for AI assistants. Its new keyboard for iOS and Android works across apps and lets users create custom AI-powered shortcuts using natural language.
shot-scraper 1.10 adds video recording capability via storyboard.yml to help coding agents demonstrate their work with Playwright automation.
MVP-Nav: RGB-only zero-shot object goal navigation framework combining semantic and physical constraints for embodied agents.
NCP-ToM benchmark evaluates LLM agents' ability to induce belief states through planning/action beyond conversation.
Engineers on the new team will embed within companies to deploy purpose-built agents, focusing on fast deployments and customer self-sufficiency.
LuckyStar 111B hybrid reasoning model from Cohere and LG CNS enables efficient multilingual tool-using agents with Korean-English support.
Comprehensive survey systematizes LLM attack surface across full lifecycle: data pipelines, agents, tools, memory, and organizational integration.
LLM agents act as constrained supervisory planners for fault recovery in process plants, validated against external safety constraints.
ACE module enables LLM agents to dynamically manage context windows by elastically retrieving discarded information, addressing trajectory length bottlenecks.
AutoTrainess enables LLM agents to autonomously conduct post-training via planning, data construction, job scheduling, and checkpoint evaluation.
FinPersona-Bench benchmark measures Mandate Salience Decay in autonomous financial agents, quantifying behavioral drift over long market horizons.
Classification framework for LLM-agent orchestration balancing autonomy, traceability, and correctness in business process management.
OKX is bringing together payments, identity and reputation into a marketplace for AI agents.