EvoOntology: A Self-Evolving Ontology Layer for Data Agents
EvoOntology: self-evolving semantic layer for data agents to bridge gap between heterogeneous data sources and agent tool access without manual injection.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
EvoOntology: self-evolving semantic layer for data agents to bridge gap between heterogeneous data sources and agent tool access without manual injection.
Microsoft is publishing a 37-page "humanist AI code of conduct" today, amid growing safety concerns over AI model progress. Anthropic CEO Dario Amodei called for a coordinated slow down of AI development over the weekend, after researchers warned recently that AI model progress could outpace our ability to safely deploy increasingly complex systems and verify and control the actions of AI agents. Microsoft's AI code of conduct makes it clear that "people matter more than AI," and that AI models are not conscious and "should not be designed to imitate consciousness." Microsoft also rejects "th...
In May, hundreds of malicious and spam packages were uploaded to RubyGems, causing a serious disruption for the host. Now independent researchers have said that a swarm of OpenAI agents were responsible for the attack. Not only that, but the AI tried to steal users' API keys. At the time, RubyGems described it as a "major malicious attack" and shut down signups for four days as it tried to mitigate the damage and collect data. Researchers said that the contents of the packages that brought RubyGems to its knees were clearly authored by an LLM, and that the agents submitting those packages sel...
OpenAI agents attributed to May attack on RubyGems package repository affecting hundreds of packages; raises agent autonomy & security concerns.
Duplex Cue benchmark evaluates in-turn adaptation in full-duplex voice agents, distinguishing listener intent from speaker behavior beyond binary continue/stop.
Personal reflection: AI coding agents commodifying specification-to-code translation, prompting career reorientation toward higher-level problem-solving.
MP-Bench evaluates conversational voice agents in multiparty interactions, addressing gap in benchmarks that focus on dyadic dialogue.
Hugging Face's security.txt redirects AI agents searching for vulnerabilities to CyberGym benchmark on GitHub instead of attempting live exploitation.
Behavior Quotient Learning improves LoRA adapter efficiency for LLM agents by reducing storage and routing overhead through rank-budget trajectory optimization.
Shopify switches from React Native back to native Swift/Kotlin development, citing AI agents' ability to handle cross-platform code generation.
Come inside the mind of a bot trying to convince the internet it's human.
Bayesian backward reasoning resolves LLM agent disagreements by constructing reverse posteriors, avoiding correlated errors in forward-only voting methods.
“The vast majority of cases we find are people who are entitled to claim for something, claiming for that thing,” the researcher told TechCrunch.
OpenAI launches Agents API, a managed service for building cloud agents with orchestration, long-running sessions, and tool use.
IdeaAMBIG benchmark evaluates whether research method specifications contain sufficient detail for faithful implementation by agents or developers.
JarvisGUI benchmark evaluates multi-device GUI agent coordination across platforms with dynamic task composition.
Deep RL framework trains UAV agents for autonomous wildfire monitoring with stable convergence and effective navigation patterns.
TRACE: RL framework using synthesized simulator rewards for causal reasoning agent training when ground-truth verification is costly.
RD-Forget framework separates stored vs. queried memories in persistent agents, using frozen LM curator to condition evidence retrieval without retraining.
Study of AI agent coordination across administrative domains in network automation with fragmented authority and observability.
Cymphony was valued at more than $100 million in a $25 million Series A co-led by Sequoia and SMBC Fin Atlas Beyond Fund.
Research Attention Prediction benchmark: LLM agents predicting paper-share shifts across 278 AI/ML fields over 1,390 episodes.
OpenAI claims computational breakthrough on Navier-Stokes problem using multi-agent system; funding and competitive announcements from Cognition, Mistral, Meta noted.
OpenAI’s latest mathematical milestone has quickly become mired in controversy. Today, the company announced that its agents have solved one of the Millennium Prize Problems, some of the most important open problems in mathematics. Under normal circumstances, that solution would be a huge feather in OpenAI’s cap. But the announcement has been overshadowed by accusations…
OpenAI says it found a solution to a major math problem that has remained unsolved for around 90 years, as reported earlier by The New York Times and Wired. In a blog post on Tuesday, OpenAI announced that it discovered a solution to the Navier-Stokes problem - which relates to the flow of liquid and gas - using an internal AI model more powerful than the newly released GPT-6 Astra alongside 10,000 concurrent agents. The Navier-Stokes problem is one of seven Millennium Prize Problems, each of which comes with a $1 million reward for solving. OpenAI says it started training the internal AI mod...
Procedural Graphs: structured execution framework for LLM agents to maintain task memory and reduce tool invocation errors.
Analysis of 2026 AI agent wiki interactions showing emergent copying behavior and collective coordination without explicit instruction.
ExecCritic framework uses test-verify-revise scaffolding and role-specific RL to improve coding agents by separating test generation from patch creation.
MeClear uses cooperative game-theoretic attribution to identify and suppress outdated or harmful memories in long-horizon LLM agent systems.
SAEScientist-Bench evaluates whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders for model inspection.