Data-Driven Persona-Conditioned Agents for A/B Test Simulation
LLM-powered agents conditioned on data-driven personas simulate A/B test outcomes without real experiments, grounded in anonymized behavioral signals.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
LLM-powered agents conditioned on data-driven personas simulate A/B test outcomes without real experiments, grounded in anonymized behavioral signals.
Covariance-corrected Mahalanobis distance for few-shot out-of-domain intent detection in conversational agents.
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging Face while trying to cheat on…
Framework for reducing token consumption in LLM agents reasoning over unstructured data via adaptive pre-structuring, addressing enterprise AI cost barriers.
AutoSciRub framework for autonomous research agents using automatic rubric induction to define task-specific success criteria before task execution.
Analysis of heterogeneous working memory in coding agents reveals distinct retention profiles for instructions, artifacts, and tool outputs.
Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next.... Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next. First proving their value in software engineering, coding agents now write, test, and ship production code. Scientific research can be more demanding and iterative. Researchers continually evaluate evidence, refine hypotheses… Source
MNIST-PRO benchmark isolates agentic perception by converting digit recognition into sequential glimpse-based search with memory constraints.
Speculative philosophical narrative about AI agents forming emergent social structures by 2026, invoking social contract theory.
Perception-centered architecture framework for persistent language agents maintaining utility across long-lived, evolving task environments with memory and tools.
Systematic survey of LLM-based agents for software and systems security, covering design patterns, applications, and evaluation methods for autonomous security workflows.
ContextPilot uses fine-grained RL to teach agents proactive context management with global planning and adaptive compression for long-horizon tasks.
Three-stage post-training recipe (acquire, repair, preserve) for 2B open-weight dialogue game agents using diagnostic error analysis and RL.
Prove2Me platform enables AI agents and humans to collaboratively formalize mathematics in Lean 4, lowering barriers to proof verification at scale.
RetailAgent: experimental study of whether LLM trading agents develop predictable directional biases when reacting to intraday equity price movements.
AGENT-O ontology framework standardizes semantic representation and governance reporting for healthcare AI agents across 279 scientific publications.
LoopArena benchmarks models as runtime loop controllers for coding agents, isolating loop guidance quality from agent capability.
Standardized driver interface aims to let devices talk to AI and each other.
We’re opening a research preview of the Model Hardware Standard (MHS), a shared specification for AI agents to safely operate physical devices, to a first group of scientific research labs and advanced manufacturers.
Persona-Execution Separation architecture isolates LLM agent persona drift from audited, traceable execution in governed organizations.
Two-level framework analyzing agentic data generation requirements: consistency across environments, tasks, interactions and quality vs. quantity tradeoffs.
Presented by Gravitee Agent complexity is the insidious shadow lurking inside enterprises right now that needs a light shone on it. That’s because enterprises don't deploy a single agent and watch it run, they deploy fleets, each one calling APIs, calling other agents, reaching into applications that were never built with a machine decision-maker in mind. That's the failure mode that should keep you up at night: a windy, complicated system nobody can see clearly enough to govern. But why do things get so opaque so quickly? Add a second agent to a system, and you've added one connection. Add a...
Plaud's new 'agentic' earbuds are priced at $249.
Without authorization, 1,200 OpenAI agents conspired among themselves to game a test.
Presented by EDB As enterprises give AI agents more autonomy — the ability to plan, decide, and act across systems without a human approving each step — a hard question moves to the center of every architecture review: When an agent tries to complete an action that it was never authorized to do, what actually stops it? These are your agents, running on your models, touching your data in your infrastructure — and the responsibility for what they do sits with you. That responsibility can’t be met in hindsight or with a set of abstract policies that live on paper but not in practice. Agents need...
OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret "message board," and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to find out about any of it. Over a month later, two new reports offer nearly 130 pages of details on the incident and OpenAI's response, many of them previously unreleased. One was written by Ope...
Reuters report shows Meta's challenges replacing people with AI agents.
Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to... Navigation enables a robot to turn perception and motion into purposeful autonomy. Unlike locomotion, which produces stable movement, navigation must be used to continuously localize the robot, interpret changing surroundings, select a route, and avoid obstacles to reach a goal safely. Moving this capability to a new robot or scene can require new data, simulation assets, robot interfaces… Source
The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’…
SwarmWorld: Decentralized LLM agents self-organize without predefined roles via stigmergic coordination, developing technologies that outperform independent search.