Gemini API Managed Agents: 3.6 Flash, hooks, and more
Google ships Gemini API Managed Agents with Gemini 3.6 Flash and tool-use hooks for production agent deployment.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Google ships Gemini API Managed Agents with Gemini 3.6 Flash and tool-use hooks for production agent deployment.
HiFi-UMI demonstrates scalable high-fidelity robot manipulation learning without real-robot anchoring by upgrading teleoperation capture quality.
Messier consolidates 957K evaluation records across 30 benchmarks and 714 agents to enable comparable cross-benchmark agent assessment.
Custom harness engineering distributes security controls for AI coding agents without vendor lock-in, enabling organizational scaling.
Speech dynamics using Lyapunov exponents and Sample Entropy identify depression biomarkers in vocal articulatory patterns.
Domain adaptation via DANN and CDAN improves CNN/transformer robustness across acoustic scene classification device shifts.
RSIBench-Data isolates LLM agents' data-centric research capability in post-training loops, decoupling research from systems optimization.
ML-based gas lift optimization using Bayesian optimization delivers 5% production uplift in unconventional oil/gas wells.
Stemma maps LLM outputs to induced decision regions to test lineage provenance, abstracting surface-form variation for reliable model genealogy detection.
The largest grid operator in the U.S. says it will cut power to large data centers to prevent blackouts starting next year.
Multi-agent LLM framework with Bayesian networks for uncertainty quantification in actuarial risk modeling and regulatory compliance.
Test-time adaptation method for LLM-based traffic forecasting on evolving sensor networks with dynamic graph structure.
Empirical analysis of attention patterns in LLM-based automated program repair to explain inconsistent patch generation performance.
kiloVAD: ultra-compact speech activity detection model using structured pruning and quantization-aware training for edge deployment.
LLM-based AI system automates discovery of quantum error-correcting codes for fault-tolerant quantum computing by iteratively optimizing code structure and decoding.
DRIFT hybrid forecasting framework combines direct and recursive action-conditioned models for ICU physiological time-series prediction with treatment interventions.
Shieldstral 3B multimodal safety classifier achieves SOTA on content moderation via unified QA framework, matching models 7x larger on text safety benchmarks.
OpenAI's Akshay Nathan details ChatGPT Work product strategy: Sites, memory, subagents, finance, no-code tools scaling from 0 to 10M users.
HiSkill hierarchical skill graph framework organizes LLM agent trajectories into structured graphs linking high-level skills to atomic operations for long-horizon tasks.
AngelSpec unified training framework adaptively switches between multi-token prediction and block-parallel diffusion for speculative LLM decoding across heterogeneous workloads.
Online learning algorithms for distributed constraint optimization applied to large-scale satellite scheduling via decomposition and iterative pricing methods.
Agentic workflow automates translation of quantum protocols to hardware experiments on Pasqal neutral-atom QPUs with researcher-in-the-loop validation.
Zero-shot sEMG movement classification via Compositional Prototype Interpolation and Synthetic Adaptation for prosthesis control without re-training per combination.
Joint agent-speculator RL aligns next tool-call prediction with deployed agent behavior by unifying speculator and agent in single model, reducing latency.
Evaluation of adversarial robustness in five Arabic language models under multi-granularity attacks exposes security vulnerabilities in non-English LLMs.
Physics-informed spectral deep operator network (SpectONet) for Euler-Bernoulli beam vibration using nonuniform sensor placement.
MC-ALFCG algorithm for stochastic nonconvex composite optimization over Markovian-sampled gradients with variance reduction.
WorkSurface-Bench evaluates enterprise agents on multi-source knowledge routing across documents, tables, and graphs with 1,151 tasks.
WALoMA: multitask 6G wireless foundation model using masked autoencoders to reduce dependence on labeled channel data.