Learning from Mistakes: Rollout-Retrieval Lifelong Policy Learning for Autonomous Driving
Proposes lifelong learning framework for autonomous driving policies to accumulate corrective knowledge from deployment errors without catastrophic forgetting.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Proposes lifelong learning framework for autonomous driving policies to accumulate corrective knowledge from deployment errors without catastrophic forgetting.
Identifies entity binding failures in tool-augmented LLM agents as distinct safety problem: selecting correct tool but acting on wrong entity.
μFlow: one-class deepfake detector trained only on real images generalizes across GAN and diffusion model generators without synthetic supervision.
ITSPACE proximal method optimizes Bures-Wasserstein objective on covariance matrices for domain adaptation and Gaussian embeddings.
TIDAL's new policy will prevent AI-generated music from making money on its service.
Knowledge distillation strategy for visual quantum RL: train classical visual teacher, freeze encoder, distill policy into classical/quantum heads.
RAPS-DA: regime-aware peer specialization framework handles heterogeneous knowledge conflicts in RAG by disentangling reliability levels of retrieved context.
Entropic Learnability Horizon (ELH) framework unifies information theory and topology to explain generalization in overparameterized deep networks.
DeepReinforce releases Ornith-1.0, MIT-licensed open-weights model (9B–397B variants) for agentic coding, built on Gemma 4 and Qwen 3.5, achieving SOTA on coding benchmarks.
Muon optimizer avoids slow saddle-to-saddle dynamics in matrix factorization; learns balanced solutions with uniform convergence rates across modes.
DR-ACI constructs prediction intervals for doubly robust pseudo-outcomes under temporal dependence via adaptive conformal inference.
Random Network Distillation enables lightweight clustering in federated learning without coupling to main training loop, reducing communication overhead.
Post-Hoc Concept Bottleneck Models often learn predictive artifacts rather than semantically meaningful concepts; proposes faithfulness evaluation beyond task accuracy.
CUDA optimization techniques (shared memory, pre-transposed weights, fused kernels) achieve 1.41x speedup on shallow networks on Tesla T4.
Multi-channel Multigrid preconditioner learned via neural networks to solve high-wavenumber Helmholtz equations with phase-space coarsening.
Google AI explains full-stack approach to AI development, covering integrated hardware, software, and model design.
A new proposal would ban the sale of Americans' health and location information to data brokers - including information people reveal to an AI chatbot like ChatGPT or Claude. In the coming weeks, Senator Elizabeth Warren (D-MA) and Representative Mary Gay Scanlon (D-PA) are planning to debut a new version of the Health and Location Data Protection Act that's better suited to the AI era. The former version of the bill, first introduced in June 2022, prohibited data brokers from collecting and selling health and location data. Four years later, it's expanded to ban other companies from selling ...
SIMAX framework generates synthetic clinician-patient dialogues with human-coded annotations to evaluate AI communication coding at scale.
Factorizable Normalizing Flows model parameter-dependent density morphing by factorizing the flow, enabling tractable learning across exponentially many configurations.
AI agents are quickly moving beyond chat. They inspect code, run tests, read documents, search knowledge bases, query internal systems, and operate for hours on... AI agents are quickly moving beyond chat. They inspect code, run tests, read documents, search knowledge bases, query internal systems, and operate for hours on behalf of a user. This unlocks productivity, but can also give agents access to sensitive enterprise data and the ability to complete tasks and take action across business systems, making a secure, governed environment essential. Source
Argues LLMs lack situated perception of physical world causality and agency; identifies this primitive as necessary prerequisite for artificial superintelligence.
COHORT automates enterprise network threat mitigation via multi-agent LLM workflow proposing, implementing, and refining mitigations on emulated topologies.
Field order in serialized structured metadata silently impacts retrieval quality; proposes permutation-invariant fine-tuning to decouple field order from embeddings.
Non-parametric causal diffusion recovery from cross-sectional steady-state data; application to gene expression without requiring longitudinal observation.
MuonSSM stabilizes state space models for long-sequence modeling via momentum pathways and Newton-Schulz transformations on input injections.
HSAP addresses cross-contamination in causal attention for packed sequences by combining sequence parallelism paradigms for hybrid-context LLM training.
Curvature-Weighted Gradient Diversity (CWGD) refines SGD convergence analysis by weighting gradient noise by inverse Hessian to better model effective optimization dynamics.
Study compares nine open-weight LLMs against human participants in networked Prisoner's Dilemma, finding selected models reproduce cooperation macro-dynamics but lack individual fidelity.
Analysis of enterprise tabular data reveals markedly different distributions and performance profiles from public benchmarks for TabPFN, TabICL, and ConTextTab models.
Three-method study across Qwen2.5-Coder-32B, Llama-3.1-8B, and Gemma-3-27B shows internal probes read situation not pre-action intent, limiting misalignment monitoring efficacy.