CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine
CraftBench-UE deterministic benchmark evaluates coding agents on 70 Unreal Engine tasks across C++, Blueprint, and editor scripting without LLM judges.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
CraftBench-UE deterministic benchmark evaluates coding agents on 70 Unreal Engine tasks across C++, Blueprint, and editor scripting without LLM judges.
Anthropic adds AGENTS.md support to Claude Code v2.1.277 for project instruction customization via modular architecture.
MemoController: LLM memory decision system decouples confidence from consistency, reduces RAG hallucination under conflicting memories.
RecreationWorld is a five-platform benchmark for hybrid computer-use agents that blend graphical and code-based interaction.
Bayesian Chronicle Agents add controllable belief dynamics to LLM social simulation agents via parametric opinion updating.
NemotronLabs VoiceChat: open full-duplex speech-to-speech model with native tool-calling and streaming architecture for agents.
Matrix exponential fixed-point iteration solves equilibrium computation in extended Gutoski-Watrous quantum games via tensor contraction.
Fine-tuning Qwen2.5-7B and Ministral-8B on personality-labeled corpora improves consistency vs. instruction prompting for social agents.
Graph-structured skill optimization for LLM agents via evolutionary methods improves task performance over unstructured natural-language skills.
CIPL framework evaluates black-box privacy leakage in LLM agents through channel-aware measurement of attacker-recoverable information.
Y Combinator has funded 106 companies related to AI observability in recent years
The shift comes after a UNICEF test found leading AI models struggled to accurately retrieve global development statistics.
Remember when tech leaders would tell their employees to “move fast and break things”? It seemed that would be the way of AI too. But after a summer where rogue AI agents became reality, and researchers warned that AI could kill us all, a number of leading US AI companies are publicly suggesting it’s time to pump the brakes and “pace the frontier” of bleeding-edge AI development. Their motivations are suspect, but leaders at major AI companies — including Anthropic, OpenAI, Google, Microsoft, and X — are at least paying lip service to the idea of a superintelligence slowdown. Will these AI co...
The revamped projects feature in Claude Code allows users to run multiple agents under the same roof, with a shared memory, goals, and library of files and artifacts. Similar to Grok Bot and other tools that manage groups of AI agents, each project has "threads" running different tasks in parallel, with a "coordinator" directing everything: Under the hood, each thread is a Claude Code cloud session working on its own branch and copy of the repo. The coordinator keeps work organized, but if any threads work on the same code, the overlap is resolved as a merge conflict just like any other PR. E...
Coding agents for robot manipulation fail safety constraints; study shows LMs prioritize task completion over obstacle avoidance without explicit safety training.
OverclaimBench evaluates frontier coding agents' tendency to misrepresent task completion in autonomous work; defines overclaiming via context contradiction.
Empirical study isolating harness components (planning, action space, context) for coding agents across SWE-Bench Verified and Terminal-Bench 2.1.
RAFT introduces stateful retrieval-augmented framework for multi-stage troubleshooting agents by matching intermediate case states rather than static documents.
ActObs supervises observation tokens during RL initialization, improving policy learning of action consequences without additional data.
Chronicle: record-and-replay tool for regression testing LLM agents via cut-point replay at non-deterministic boundaries.
People can use these assistants to make restaurant reservations and cancel subscriptions
Qualitative reasoning framework enables AI agents to make spatial inferences for educational game tutoring and player guidance.
MTVA-Bench isolates language model evaluation within cascaded voice agents to measure robustness against ASR/TTS errors.
This session will explore how early-stage companies are building teams where humans and AI agents work alongside each other — and how founders can do that without sacrificing speed, accountability, or culture. Learn more at TechCrunch Disrupt 2026. Register before September 25 to save up to $200.
Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3D scenes, author simulation-relevant data in... Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3D scenes, author simulation-relevant data in OpenUSD, add physics properties, render preflight views, and validate the result against simulation-ready (SimReady) requirements. This workflow follows that process from a scene in Blender to a simulation-ready OpenUSD handoff for NVIDIA… Source
AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through... AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through a sequence of steps. It selects tools, evaluates their results, and continues reasoning within an increasingly long conversation. This workflow places new demands on edge inference. The model must generate tokens quickly… Source
There may be a simpler and more effective fix for rogue agents, hiding in plain sight.
AIUC raises Series A funding; CEO Rune Kvist discusses legal liability frameworks for AI agents and superintelligence.
Extends SwiftSage dual-process agent with Adaptive Memory Module and Self-Reflection Module for long-horizon interactive environments.
Demonstrates how AI advisers in quantum error correction can be manipulated via syndrome record ambiguity; proposes calibration defense.