CiteGuard-RAG: A Validation-Centered AI System for Evidence-Grounded Question Answering
CiteGuard-RAG adds citation validation, grounding checks, and selective regeneration to RAG systems to ensure evidence-grounded answers with explicit refusals.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
CiteGuard-RAG adds citation validation, grounding checks, and selective regeneration to RAG systems to ensure evidence-grounded answers with explicit refusals.
The code of conduct lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles.
Theoretical analysis refutes isotropy assumption in spectral contrastive learning; shows task priors shape optimal singular subspaces for frozen representation transfer.
AlgoEvo enables autonomous agents to dynamically inspect, edit, and improve code via runtime feedback and a paradigm-agnostic skill hub for algorithm discovery.
Atria Dawn is a foundation agentic LLM trained via Verifiable Experience Pipeline for scientific research workflows, evaluated on 16 real-world benchmarks.
As the complexity of High Performance Computing (HPC) ecosys- tems continually increases, achieving optimal performance becomes a challenge. Traditional performance autotuning techniques pro- vide promising means to navigate this complexity, these techniques remain computationally intensive and require many evaluations to find optimal configurations. This work proposes an autotuning framework that designs a machine learning-based ensemble LLVM Intermediate Representa- tion (IR) ranker, Neural Configuration Scorer (NCS). NCS ranks the performance of IRs sampled by a transfer-learning-based aut...
Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions. If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed. But requiring human approval at every step makes attention a bottleneck. Delegating review to other AI agents raises the same alignment problem: the reviewers may themselves be misaligned. We identify a condition on a reviewing panel that is weaker than individual alignment yet necessary and sufficient for a guara...
We introduce loss-conditioned state execution, a model-agnostic method that decides whether to execute a world model's fixed feasible proposal or retain the current state. Predictive informativeness alone, however, does not establish whether an update will reduce downstream loss. Occurrence ranking can approach perfection while persistence remains the unique absolute-loss Bayes action. Two transition laws can also share occurrence information and conditional variance yet require opposite absolute-loss decisions. We formalize state movability as the existence of a loss-reducing feasible correc...
Visual Retrieval-Augmented Generation (VRAG) empowers models to navigate and answer queries about visually rich documents by retrieving relevant page images as visual evidence and reasoning over their content. However, effectively utilizing this visual evidence is usually impeded by two main challenges. First, answer-relevant evidence is sparse and may be concentrated in a small region of one page or dispersed across multiple pages. Second, existing agentic methods often generate answers based on raw exploration trajectories or compressed textual memories rather than an explicitly organized s...
KnowBench introduces Effort Reduction metric for evaluating clinical AI systems based on clinician acceptance of auto-generated work products rather than reference similarity.
Study of density ratio estimation for importance-weighted regression under target shift with continuous outputs, providing finite-sample convergence rate analysis.
Richard Socher launches Recursive, an RSI-focused startup valued at $5B, backed by his NLP expertise and You.com CEO experience.
Backdoor attack vulnerability in pretrained world models reused for downstream control tasks, enabling hijacking via poisoned checkpoint without explicit triggers.
Google DevFest 2026 community event series launches with 800+ global meetups focused on agentic AI development.
A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in…
MoveBench: large-scale benchmark for probabilistic wildlife movement forecasting with 2.6M GPS locations across 110 species and 1.6B environmental rasters.
EvoOntology: self-evolving semantic layer for data agents to bridge gap between heterogeneous data sources and agent tool access without manual injection.
Transfer learning approach using earth observation data to estimate socioeconomic conditions for forcibly displaced populations between survey rounds.
Spiking neural network framework for cyber threat detection handling categorical identifiers and temporal patterns in asynchronous heterogeneous event streams.
Sylvas: device scheduling algorithm for federated continual learning quantifying edge device contribution to global model performance under resource constraints.
Lightweight streaming ASR head added to full-duplex speech-to-speech models enabling simultaneous listening/speaking with user transcription capability.
Sequential Adapter Stacking transfers knowledge from high-resource to low-resource languages in Whisper multilingual ASR via parameter-efficient adapters.
DescaPE decoding framework uses internal model signals to suppress hallucination during LLM generation via layer-wise factual salient detection.
Deep learning system for financial credit risk early warning integrating heterogeneous data via neural networks and attention mechanisms.
Model merging approach integrates external language models into ASR systems to reduce computational overhead vs. shallow fusion at inference.
Multimodal foundation model for zero-shot brain signal analysis using language guidance to address representational misalignment in neuroscience.
Empirical study evaluating LLMs as financial user simulators via controlled paper-trading with 120 participants and virtual funds.
Bench2Dex benchmark for visuo-tactile bimanual dexterous manipulation across 12 simulated dexterous hands with standardized tactile sensor evaluation.
Multi-block variance reduction method for finite-sum coupled compositional optimization problems with stochastic gradient estimation.
DIST Pyramid framework integrates data storytelling with interpretable ML to explain AI decisions to non-experts without revealing model details.