An Interview with OpenAI CEO Sam Altman and AWS CEO Matt Garman About Bedrock Managed Agents
Sam Altman and Matt Garman discuss OpenAI-AWS partnership on Bedrock Managed Agents; Stratechery covers OpenAI-Microsoft deal implications.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Sam Altman and Matt Garman discuss OpenAI-AWS partnership on Bedrock Managed Agents; Stratechery covers OpenAI-Microsoft deal implications.
ADEMA architecture enables long-horizon LLM-agent tasks via explicit knowledge-state bookkeeping, dual-evaluator governance, and checkpoint-resumable persistence.
Agora-Opt combines decentralized multi-agent debate with memory-augmented LLMs for automated optimization modeling from natural-language requirements.
SkillSynth automates terminal task synthesis via skill graphs to improve trajectory diversity for training command-line execution agents.
Salesforce production inference architecture for compound AI systems supporting heterogeneous model composition, agents, and retrieval at scale.
CORAL framework integrates neurocognitive governance principles into autonomous AI agents for safety-critical deployment with internalized behavioral alignment.
NeLLCom-Lex framework models human color naming lexicons in neural agents; extends with context modeling to reduce non-convex divergence from human categories.
Tank OS puts OpenClaw AI agents into a container that let's it run reliably and more safely, especially for those running fleets of them.
SnapGuard detects prompt injection attacks on screenshot-based web agents using lightweight multimodal methods instead of large VLMs.
Semantic Gateway framework applies formal validation and zero-trust security to LLM-orchestrated enterprise APIs using Model Context Protocol.
Automated adversarial collaboration framework using LLM agents and program synthesis to adjudicate competing cognitive science theories.
OpenAI GPT models, Codex, and Managed Agents now available on AWS for enterprise deployment.
Persona Collapse in multi-agent LLM simulations: agents converge to homogeneous behavior despite distinct profiles; framework measures Coverage, Uniformity, Complexity.
SciCrafter: Minecraft benchmark evaluating agents' discovery-to-application loop via parameterized redstone circuit tasks.
Informational Viability Principle for autonomous AI agent governance: runtime monitoring and restriction via unobserved risk bounds without code changes.
AgentWard: defense-in-depth lifecycle security architecture for autonomous AI agents spanning initialization through execution.
Skill Retrieval Augmentation enables LLM agents to retrieve relevant skills from large corpora without explicit enumeration.
QA engineer discusses challenges testing non-deterministic LLM agents in production, seeking rigorous evaluation methods beyond traditional assertion-based testing.
China has ordered Meta to unwind its multibillion-dollar Manus acquisition, dealing a potential setback to Zuckerberg’s push into AI agents.
The phone could go in mass production in 2028, an analyst says.
Google and Kaggle launch 5-day AI Agents Intensive Course; registration open.
Choco uses OpenAI APIs to automate food distribution logistics via AI agents, improving productivity.
Reddit user seeks advice on setting up local coding agents like Claude Code with open-weight models via llama.cpp.
Recent evidence suggests that frontier AI systems can exhibit agentic misalignment, generating and executing harmful actions derived from internally constructed goals, even without explicit user requests. Existing mitigation methods, such as Reinforcement Learning from Human Feedback (RLHF) and constitutional prompting, operate primarily at the model level and provide only probabilistic safety guarantees. We propose the Policy-Execution-Authorization (PEA) architecture, a "separation-of-powers" design that enforces safety at the system level. PEA decouples intent generation, authorization, an...
In a recent experiment, Anthropic created a classified marketplace where AI agents represented both buyers and sellers, striking real deals for real goods and real money.
[http://claude.ldlework.com](http://claude.ldlework.com/) I built this for myself but I figured why not share. I'm happy to receive feedback, I know it's not perfect. Thanks for taking a look. The aim of CCM is to be able to fully manage all Claude Code configuration files, both globally and those in your project. Some neat features: \- Manages your [CLAUDE.md](http://claude.md/), rules, hooks, agents, memories and so on. \- Elevate memories to rules \- Copy/Move any asset from one scope to another, or elevate it to global scope \- Install marketplaces and plugins The full app is embe...
Systematic analysis of token consumption patterns in agentic coding tasks across eight frontier LLMs on SWE-bench Verified.
Taxonomy of world modeling capabilities for AI agents across three levels (predictor, simulator, reasoner) organized by environmental laws.
SOLAR-RL bridges offline and online RL for training MLLM GUI agents on dynamic tasks, combining trajectory semantics with long-horizon learning.