Plover: Steering GUI Agents through Plan-Centric Interaction
Plover externalizes planning in vision-based GUI agents, enabling user inspection and correction of task plans for autonomous interface automation.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Plover externalizes planning in vision-based GUI agents, enabling user inspection and correction of task plans for autonomous interface automation.
Critical analysis of item response theory reliability for AI benchmarks, highlighting regime mismatches between IRT assumptions and benchmark data distributions.
Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap —...
Hybrid neural-physics ODE framework with RTS smoothing for state/parameter estimation in dynamical systems with unknown components.
T²MLR fuses cached middle-layer representations across decoding steps to enable persistent reasoning in transformers with minimal inference overhead.
Benchmark evaluates six MLLMs on scientific visualization literacy using 49 standardized assessment items across 8 visualization techniques.
Study investigates linear representations of grammaticality in neural language models beyond probability-based measures.
MedFailBench is clinician-built open-source benchmark categorizing medical AI failures by severity and safety gate type with 44 synthetic cases.
Opinion piece examines AI's transition from tool to autonomous research participant via industrialization of science, citing DOE Genesis Mission.
Research investigates scaling laws for Behavior Foundation Models in humanoid robot control using large-scale behavioral data.
On-Policy Delta Distillation introduces delta-signal reward for RL post-training, measuring difference between teacher and base models.
A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and... A video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and applications to be useful. These include content management systems, messaging platforms, databases, ticket queue, and escalation paths. This integration is challenging because video systems, enterprise knowledge bases… Source
Google integrates third-party app connections into Search's AI Mode for secure data access and interaction.
With this new update, Google is expanding AI Mode beyond answering questions and into completing tasks across the apps they use regularly.
Google is giving its AI note-taking app a new name. The company announced on Thursday that NotebookLM is becoming Gemini Notebook, but will remain a standalone app even as it integrates more deeply across Gemini and Google Search. Google first revealed Gemini Notebook - then called Project Tailwind - in May 2023 before widely releasing the app just months later. Over the past few years, Google has been adding new features to the app to help organize and make sense of your notes, such as the ability to summarize them as AI podcasts, narrated slideshows, and TikTok-style clips. It recently star...
Google Vids adds Gemini Omni support and personal avatar features for video generation and editing.
OpenAI introduces age-appropriate safeguards, parental controls, and learning tools for ChatGPT teenage users.
Agentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks, storage... Agentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks, storage accesses, and network transfers before a final answer is produced. As more agents run at once and carry context across steps, users, tools, services, and sessions, infrastructure must move, protect, retrieve, and reuse data fast enough to keep… Source
Large-scale political bias audit compares Grokipedia (Grok-written encyclopedia) and Wikipedia across 1,394 article pairs on neutrality.
Companies coming to market are raising money at fastest pace this century.
Diagnostic study isolates and evaluates five visual world models (DreamerV3, DIAMOND, TWISTER, Simulus, STORM) in Atari Pong.
DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place orders from the terminal, marking another step toward software designed for AI agents instead of just humans.
Thinking Machines Lab releases Inkling, a 975B-parameter open-weights MoE multimodal model trained on 45T tokens.
NIFA extends FPGA-integrated ReRAM in-memory computing to support nonlinear operations for efficient ML inference.
You may have heard that OpenAI released its first piece of hardware this week. You may not have heard about the ChatGPT basketball.
Categorical framework (LINCS) addresses non-compositionality in ML via tangent category sketches and universal factorization.
Hierarchical Global Attention + tiered KV storage enables 16K-token fine-tuning on 16GB VRAM, 8× longer than dense attention baseline.
Multi-agent LLM framework using SFT+DPO enables sustained partisan behavior in political coalition simulation, circumventing RLHF neutrality bias.
AlphaWiSE: post-hoc weight interpolation maintains cross-modal alignment in CLIP during continual multimodal learning via per-tensor scalar coefficients.