Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems
Semantic Reward Collapse explains miscalibration in RLHF systems, linking preference optimization to performative certainty and hallucinations.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Semantic Reward Collapse explains miscalibration in RLHF systems, linking preference optimization to performative certainty and hallucinations.
Sam Altman testifies that Elon Musk sought total control of OpenAI to pass to his heirs; governance dispute.
Google unveiled its new AI-first Googlebooks laptops, more agentic Gemini features, vibe-coded Android widgets, Gemini in Chrome, refreshed Android Auto, and more ahead of I/O.
OGLS-SD improves LLM reasoning via on-policy self-distillation with outcome-guided logit steering to correct reflection bias.
Google has revealed its vision for the AI laptop of tomorrow.
Google has big plans for Android in 2026, and most of it is AI.
Q-DAPS estimates question difficulty for LLM QA evaluation using entropy of answer plausibility scores.
AI-generated widgets are among the features coming to Android this year. | Screenshot: Google Would it shock you to hear that Android 17 is filled with new AI-enabled features, like improved dictation and vibe-coded widgets? Fortunately, that's not all. The platform is getting non-AI updates too, from an emoji overhaul to a new screentime tool that helps you avoid distracting apps. Google has just revealed the biggest changes coming in its next OS update as part of its dedicated Android Show, ahead of next week's big I/O developer conference. The Android software updates came alongside a teas...
Gemini Intelligence comes with a Liquid Glass-ish visual treatment. | Image: Google It is, once again, Gemini season. Google is announcing a host of new Gemini features during its pre-I/O Android showcase, many of which aim to help use your phone for you. You'll find Gemini in more places, like Chrome on Android, in your autofill suggestions, and all up in your apps - if you want. Google also has a new name for us to remember, because it just can't help itself: Gemini Intelligence. It "brings the very best of Gemini to our most advanced Android devices," according to Google's director of Andr...
Gemini Intelligence will also include Gboard based dictation and form filling capabilities
As the AI legal services industry heats up, Anthropic is launching its own suite of features designed to assist law firms.
The new feature will first launch on the latest Samsung Galaxy and Google Pixel phones this summer.
Google's transcription feature will initially launch with Samsung Galaxy and Google Pixel phones
Level-playing-field evaluation framework compares controlled text generation systems across datasets with uniform methodology.
Random Matrix Theory detects overfitting onset in neural networks via Correlation Traps without accessing train/test data.
Deep learning method (W-Net) for detecting asteroids in TESS time-series data with rotation-invariant training.
Graph-based representation learning framework for segmenting sparse structures in large-scale images.
Multi-agent RL framework using events to trigger dynamic behavioral transitions beyond fixed role assignments.
Semi-supervised confidence detection in speech using Whisper encoder with uncertainty-aware pseudo-labeling.
TokenHD: token-level hallucination detection pipeline for reasoning tasks with improved granularity over step-based methods.
Analysis of popularity bias in OLMo LLMs traced to pretraining exposure in Dolma corpus via entity-level statistics.
Adaptive policy optimization removing hyperparameter tuning for RL post-training under distribution mismatch.
DRIFT: discrete flow matching for offline-to-online RL using CTMC policies with advantage weighting.
ProfiliTable: multi-agent framework using profiling-driven agentic workflows for table cleaning and transformation.
LLM agent framework for post-hoc crop yield forecast correction using domain tools on strawberry and corn data.
GAP framework fixes feature-space mismatch in multimodal LLM visual reasoning by aligning latent token generation with input embedding norms.
Study shows passage convergence—how effectively hints eliminate wrong answers—improves LLM performance on inferential QA over retrieved answers.
MetaColloc meta-learns neural basis functions offline to solve PDEs at test time without retraining, replacing optimization with collocation assembly.
The feature is designed to help people get real-time context about trends and breaking stories, as well as receive recommendations, all within conversations.
Frontier models (Opus 4.6, GPT 5.4, Gemini 3.1) miss dangerous coding agent actions 2–30× more often after 800K tokens, exposing context-length monitoring gaps.