Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork
Ad-hoc teamwork extension for multi-task scenarios where agent partner capabilities are hidden; frames collaboration as joint planning problem.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Ad-hoc teamwork extension for multi-task scenarios where agent partner capabilities are hidden; frames collaboration as joint planning problem.
e-Commerce search system using intent-conditioned recall expansion to improve discovery of substitute and complementary items.
ML-based repair ranker for continuous algebraic constraint systems outperforms random enumeration on structural augmentation selection.
SpecFirst framework improves LLM agent program synthesis from scratch by separating behavioral specification elicitation from code synthesis.
OmegaUse-OfficeVal benchmark evaluates LLM agents on 100 long-horizon office-suite tasks with cost grounding and human labor baselines.
Anatomy Contextualized Adaptation framework fine-tunes CT vision-language models with anatomy-specific signals via lightweight adaptation.
WindCastNet uses satellite scatterometer data for nowcasting offshore wind, improving on traditional NWP for intraday forecasting.
MindForge trains small LMs on whole-lifecycle program synthesis from scratch via source-free generation across full software development cycle.
Cost-sensitive conformal prediction benchmark addresses minority-class under-coverage in imbalanced high-stakes decision systems across six domains.
Hardware framework integrating memristor-based reservoir computing with CMOS logic for branch prediction in pipelined CPU cores.
DLAM learns distributional latent actions from action-free video with temporal constraints for vision-language-action robot models.
LLM-assisted writing may reduce linguistic diversity at population scale through repeated coevolution between authors and shared models.
Theoretical characterization of minimal Markov sufficient statistics for holonomy-cover POMDPs via stable quotient abstraction.
AgentMap: multi-agent LLM framework for unified ontology matching that discovers both equivalence and subsumption relations.
Voronoi histogram-based vectorization method for Expected Persistence Diagrams as alternative to functional approximations.
Deploying an AI coding assistant in a regulated, sovereign, or source-sensitive environment, often comes with challenges. Three common issues are: the source... Deploying an AI coding assistant in a regulated, sovereign, or source-sensitive environment, often comes with challenges. Three common issues are: the source cannot leave the network, the assistant occasionally invents package names that introduce supply-chain risk, and there is no audit trail when a generated change ships a defect. This tutorial walks you through how to self-host a validated… Source
MMAC: 5,638-clip benchmark for audio captioning across 15 evaluation dimensions, targeting fine-grained free-form descriptions for AudioLLMs.
Hierarchical spatio-temporal transformer for multi-level emergency department demand forecasting; healthcare application, not core AI frontier.
Critical-transitions approach to seizure detection; medical AI application with limited relevance to frontier LLM/agent research.
Analysis of ~100B language models' decodable night-sky representations in residual streams; mechanistic interpretability finding with 65-85% variance captured.
InferScale: GPU-native KV injection system for personalized LLM serving with persistent memory, reducing TTFT overhead from repeated prefill.
SciFigQual-Bench: benchmark for scientific figure quality assessment using full-manuscript context; evaluates caption alignment and visual misleadingness.
Cost-aware stopping mechanism (CAM-DF) for LLM agent tool acquisition balancing task coverage against cost, context load, and privacy.
On-Policy Distillation defense against fine-tuning poisoning attacks; routing-based approach to template-robust LLM safety realignment.
MemSecBench: benchmark tracing malicious content lifecycle in agent memory systems from persistence through recall to repair across backends.
Field codes for distributed optimal transport coupling sampling; mathematical framework for empirical Monge maps with certified marginals.
Google DeepMind releases Lyria 3.5 in Google Flow Music with improved musicality, lyrics, vocals, and creative control.
Parallel Trajectory Tempering improves Energy-Based Model training stability on multimodal scientific data via better MCMC mixing.
Single-beat cuffless blood pressure estimation via ear-PPG and ECG with lightweight hybrid learning for wearables.
HT-PAder: parameter-free online convex optimization algorithm for non-stationary environments with heavy-tailed noise.