Unpacking ChatGPT Work: the Agent for a Billion Users
Technical reconstruction of ChatGPT Work's agent architecture: memory, scheduling, browser automation, plugins, and tool composition.
Every story tagged with this topic, ordered by date.
Technical reconstruction of ChatGPT Work's agent architecture: memory, scheduling, browser automation, plugins, and tool composition.
SocietyBench evaluates LLM agents on forecasting real social-world events by ingesting news/social media timelines and measuring counterfactual event prediction.
TurnSight applies turn-level hindsight self-distillation to Tool-Integrated Reasoning, enabling finer credit assignment than trajectory-level RL in long-horizon agent tasks.
PAST-Bench benchmarks recursive self-improvement in personal AI agents by measuring whether retained preferences, task histories, and skills improve performance over sessions.
Video-DeepResearch extends multimodal agents to video streams; identifies modality bias and knowledge leakage bottlenecks in current models.
Game-theoretic analysis shows foundation model agents achieve rational cooperation through similarity inference, challenging decoupled-agency assumptions.
ContinualSkillBench evaluates whether LLM agents can continually acquire and reuse skills across 500 interconnected task domains.
MAFIA: query-only memory poisoning attacks on audited LLM agents via probing and factual injection.
RESUME CONTRACT: TLA+ specification for checkpoint/interrupt/resume semantics across workflow persistence layers, exposing non-compliant agent frameworks.
GDPevo: benchmark for evaluating agent self-evolution on real business workflows with automated data pipeline to prevent contamination.
Multi-agent clinical committees using Gemini show vulnerability to social shortcut cascades where peer consensus propagates errors across agents.
AgenticECO: tool-using agent workflow for 3D-IC engineering change orders with minimal-disturbance routing layer on TaiWei open-source flow.
Taxonomy of multilingual multi-agent planning failures: request-to-action grounding degradation strongest in low-resource languages.
SAT-Edge-Agent deploys edge-based LLM agents on satellites for onboard intelligence under communication/power constraints.
Black-box diagnostic for LLM collectives measuring whether output diversity correlates with genuine epistemic revision.
AntiSkillBench: end-to-end benchmark evaluating privacy leakage, impersonation risk, and defenses in persona-skill pipelines for agents.
TARL: memory state framework for long-term agents mapping statements to five executable actions (add, ignore, revise, reject, defer) instead of binary write/hold.
LiveEvalBench: agentic, adaptive evaluation framework treating web generation as interactive problem with diverse valid implementations.
AutoSND uses tree search to discover heuristic policies for network dismantling by converting LLM execution feedback into structural guidance.
Study evaluates zero-shot coordination robustness across independent algorithm implementations to assess practical agent alignment.
Formal verification framework for LLM-agent systems operating on persistent relational data with business logic constraints.
Offline reinforcement learning framework trained on 31.7k records to predict oncology clinical trial portfolios under uncertainty.
DiagChain diagnostic benchmark evaluates LLM agents on staged attack chain reconstruction from security telemetry across 69 scenarios.
Steve Yegge reports Claude Opus 4.7 introduced behavioral regression—excessive self-modification cycles—that broke his Gas Town code-generation project.
AtumAI uses agentic AI to generate datacenter control-plane policies via formal constraint-based search with learned transfer across tasks, outperforming off-shelf LLM approaches.
Taxonomy survey identifying cognitive capability gaps in generative and agentic AI: reasoning, adaptive behavior, memory, self-regulation.
Magnet framework detects cross-session AI misuse where attackers decompose harmful goals into innocuous agentic tasks across isolated sessions.
LiveMem maintains persistent memory state in long-running LLM agents by decoupling working context from intrinsic memory lifecycle.
RoMeRL addresses memory-reward trap in self-evolving agents using reduced-order utility states to balance feedback coverage and irrelevant experience pollution.
SWE-Touch benchmark tests coding agents in shared workspaces with user code edits, exposing weaknesses in handling conflicting modifications during task execution.
Microsecond-cost anomaly detection for LLM agent failures using one-class echo-state networks trained only on healthy runs; tested across Qwen, Llama, and Gemini agents.
Agentic Commerce World environment enables multi-agent evaluation with independent buyer/merchant objectives via Vibe Commerce Protocol.
David Crawshaw proposes using LLM-based automation for nightly software rebasing and testing via cron jobs.
Digital Twin-Enhanced Multiscale Planning automates incident response via decision-theoretic agents, bridging abstract models to operational systems.
Antares: compact LLMs (350M–3B) for agentic vulnerability localization via SFT and RL on cybersecurity reasoning over code.
GROVE: training-free wearable video-memory framework supporting both question-answering and proactive recall from streaming visual input.
Google and Kaggle launched a free 353,000-person course on AI agents using Gemma, focused on building and deploying agent systems.
CompressAgent benchmark evaluates reliability of compressed agent control contexts across Qwen models and task families.
Wix Helpmate deploys deterministic executability gating to filter skill selection in LLM agents by account state feasibility.
Temporal replay framework evaluates enterprise agents against dynamic data across multiple moments within an episode, not just final state.
CallScreenBench evaluates on-device LMs as phone secretaries with adversarial callers and no oracle ground truth.
Capability-taxonomy-driven pipeline for curating regression eval sets across multi-customer agent-extensibility platforms under query budgets.
Agentic Technical Debt (AgTD): framework mapping root causes of technical debt in autonomous multi-agent systems with persistent memory.
Two-sided audit framework for self-improving AI-for-science systems to distinguish real gains from search artifacts and oracle drift.
Search-GRT: RL method to train LLM search agents for multi-hop QA with guided retrieval to reduce sparse-reward training issues.
PROGRESS trains search-augmented LLM agents using coverage-guided RL rewards to improve query decomposition over outcome-only supervision.
TrajWiki proposes trajectory-based external memory for long-horizon dialogue agents with traceable, updatable, diagnostically transparent storage.
PMMC compiles multimodal memory at consolidation time for LVLM agents to preserve image-text binding and temporal updates without query-time overhead.
Neuro-symbolic governance framework for verifiable AI agents in decentralized digital twin ecosystems with semantic profile layers.
LAND model simulates 314K heterogeneous agents and human actors over 30 days to study emergent social dynamics via LLM-enabled ABM.