Uncertainty Is Not Enough: Value-of-Information Routing for Mixtures of LoRA Experts
VI-MoLE routes inputs through LoRA expert mixtures using value-of-information rather than uncertainty, avoiding spurious expert activation.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
VI-MoLE routes inputs through LoRA expert mixtures using value-of-information rather than uncertainty, avoiding spurious expert activation.
MedPRESS: 600-dialogue benchmark measuring patient-pressure-induced sycophancy in LLMs across medication, self-care, and triage scenarios.
Moment closure enables distribution-aware planning in stochastic RL without restrictive policy assumptions, addressing target variance and covariance gaps.
Magnet framework detects cross-session AI misuse where attackers decompose harmful goals into innocuous agentic tasks across isolated sessions.
LiveMem maintains persistent memory state in long-running LLM agents by decoupling working context from intrinsic memory lifecycle.
UMDP optimization balances performance and operational constraints by selecting minimal policy sets robust across uncertain environment models.
RoMeRL addresses memory-reward trap in self-evolving agents using reduced-order utility states to balance feedback coverage and irrelevant experience pollution.
Finite sample characterization of log-likelihood ratio statistics in logistic regression with nonasymptotic bounds independent of design regularity.
Identity abduction—inferring two structures as one object via representational grounding—occurs without continuous embodiment in scientific hypothesis generation.
CMuon optimizer stabilizes Diffusion Transformer training by chunking momentum orthogonalization on fused tensor weights, improving late-stage convergence.
SWE-Touch benchmark tests coding agents in shared workspaces with user code edits, exposing weaknesses in handling conflicting modifications during task execution.
DyFrDet detector improves small object detection via dynamic frequency suppression and label disambiguation to handle visual cues scarcity.
Study proposes longitudinal measurement framework to detect cognitive, developmental, and socio-affective behavioral changes from long-term LLM interaction rather than short-term evaluations.
Computational and statistical analysis of c-rectified flow framework used in FLUX.1 and Stable Diffusion 3, proving guarantees for cost-aware velocity field projection.
Analysis of 18 open-source LLMs showing cultural bias in mythology knowledge; models encode cross-cultural distinctions in residual streams but fail to decode non-Western traditions.
Private Bayesian bootstrap technique using group-level blocking weights to privatize both point estimates and uncertainty quantification for statistical reporting.
House spending records show OpenAI's ChatGPT dominates paid AI use on Capitol Hill, with congressional offices relying on the chatbot to draft memos, summarize legislation, and assist constituent communications.
CTRAG framework uses in-context retrieval with LLMs for automated regulatory compliance checking across financial, privacy, and cybersecurity domains.
DiffeoAfford derives surgical attention supervision from tissue tracking and instrument trajectories to enable real-time anticipatory framing during laparoscopy.
Study shows Qwen-VL achieves 87.3% accuracy on vehicle damage classification but fails spatial grounding for fine-grained defects; proposes dedicated segmentation layer.
Microsecond-cost anomaly detection for LLM agent failures using one-class echo-state networks trained only on healthy runs; tested across Qwen, Llama, and Gemini agents.
Cross-modal analysis of scientific formulae shows weak correspondence between syntactic and semantic representations despite joint modeling improving retrieval.
Aggregate-then-Calibrate framework combines heterogeneous human judgments with model scores to assess human-centered tasks lacking ground truth with theoretical guarantees.
LTL-to-LTLf+ translation reduces computational complexity of temporal logic specifications for RL and planning tasks.
Pinterest deploys vision-language models for automated relevance evaluation in search, reducing human annotation cost at scale.
ParEvalLayer detects biased partial evaluations of LLM agents, preventing premature benchmark conclusions from incomplete task runs.
Solution Hacking identifies shortcuts where LLMs achieve correct answers without valid reasoning on frontier science benchmarks.
Agentic Commerce World environment enables multi-agent evaluation with independent buyer/merchant objectives via Vibe Commerce Protocol.
POMDP framework using active inference separates opponent intent from execution noise in noisy multi-agent social dilemmas.
xPress improves speculative decoding by refining block-diffusion drafters through parallel refinement of conditionally dependent tokens.