Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
Proposal for agent serving systems to read actual tool duration from running tools rather than pre-call estimates, reducing KV cache memory waste.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Proposal for agent serving systems to read actual tool duration from running tools rather than pre-call estimates, reducing KV cache memory waste.
ReFigBench evaluates multimodal agents on reconstructing scientific figures as editable PowerPoint, isolating visual perception from planning and tool harness failures.
Infinite-parameter LLMs adapt weights from live interaction data via MoE, enabling models to learn facts and corrections post-deployment without retraining.
Interpretable multi-instance learning predicts AML mutations from flow cytometry hours faster than standard molecular testing using decision-tree classifiers.
Causal analysis of VLM attention heads reveals general-purpose semantic features enabling OCR, with heads outputting interpretable tokens across diverse image regions.
Compositional Policy Violations expose governance gaps in agentic workflows where step-level compliance passes but composed execution violates organizational policies.
WaveTLM separates task-object reliability from predictive quality in time-series LMs via ExecTS-QA benchmark spanning forecasting, imputation, classification, anomaly detection.
ProgramDistill benchmarks coding agents on inferring and implementing features from interactive reference web apps via mine-craft-patch factorization pipeline.
Contextual embeddings from SciBERT outperform frequency analysis for tracking semantic meaning shifts in domain-specific scientific terminology across 2010–2024.
Convergence framework for deep V-learning decomposes update error into six residuals under concentrability assumptions, establishing sharp action-gap bounds.
CERA-MoA co-evolves routing and agent policies via reinforcement learning, enabling Mixture-of-Agents to specialize dynamically as agent capabilities improve.
The robotics industry is still waiting for their breakthrough into day-to-day life. Nvidia's Les Karpas has an answer as to why at TechCrunch Disrupt 2026. Register before September 25 to save up to $200 on your pass.
Zero-shot cross-lingual handshape recognition transfers ASL phonological features to Catalan Sign Language via decomposed feature prediction.
Production QA system for normative documents handles version control, jurisdiction scope, and source traceability at 73K-query scale.
FRAUDSkill applies frozen-weight structured optimization to audio-language models for adaptive telecom fraud detection without retraining.
Analyzes structural stability of graph-aware Schrödinger bridge generative models under graph perturbations using Wasserstein distance.
Security researcher Matthew "Zigula" Gore-Kormanik was analyzing a fraudulent dating app called Dora when he got a pop-up message saying he was receiving a call from Jennifer. According to her bio, she's a 41-year-old Sagittarius with red hair, blue eyes, and piercings. She likes music, horror movies, nightlife, and sports. Gore-Kormanik answered the call, but didn't see Jennifer in his video feed. He saw a tapestry that was moving, probably due to a fan, and heard weird distortion in the background. After the call ended, "Jennifer" messaged him to say she'd had fun and "your voice is way bet...
TeleAntiFraud 2.0 benchmark refreshes monthly with evolved scam patterns, distinguishing fraud from legitimate near-domain audio via Mixed-Tree pipeline.
Replicates Edit Flows and EvoFlows for antibody optimization, revealing both use identical jump-process mechanics for variable-length edit generation.
Multi-step NER annotation correction framework uses self-training and dual-threshold inference to reduce label noise in low-resource datasets.
Evaluates multi-agent tool-augmented deductive reasoning via Clue board game environment, testing LLM consistency across extended reasoning chains.
Empirical evaluation of LLM ability to translate informal natural language goals to PDDL for automated planning in video game testing.
Audit of ChatGPT, Google Gemini, and Google Search AI product recommendations via 2,528 consumer queries reveals bias patterns across platforms.
Threads is rolling out new tools for podcasters, including profile cards, episode links, transcripts, guest tags, posting reminders and audience insights, as Meta looks to make the X rival a bigger hub for podcast promotion and discussion
Meet the next five top-tier investors judging the Startup Battlefield 200 contenders live at TechCrunch Disrupt 2026. Register before September 25 to save up to $200 and don't miss a moment of the ultimate startup pitch competition.
Sign-based variance reduction methods achieve optimal convergence rates in distributed heterogeneous settings via server-side gradient tracking.
QFWP-ANO uses classical hypernetworks to dynamically program quantum neural circuit parameters and multi-qubit measurements for time-series forecasting.
DyMT-ESB evaluates social bias in multi-turn LLM interactions using response-conditioned dynamic dialogue to measure stereotyping harms.
Fallacy detection benchmarks conflate scheme recognition with fallacy detection; negative class design inflates false-positive rates and masks poor generalization.
STRETCH framework uses dynamic difficulty adjustment and cognitive scaffolding to enable continuous LLM self-improvement beyond capability stagnation.