Optimal No-Regret Learning for Repeated Prophet Inequality
Efficient algorithm for repeated prophet inequality under prefix feedback achieves Õ(√T) expected regret matching lower bounds.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Efficient algorithm for repeated prophet inequality under prefix feedback achieves Õ(√T) expected regret matching lower bounds.
Statistical framework tests whether LLM-as-judge peer-review metrics capture substantive quality or merely surface-level linguistic polish.
Kernel-based mechanistic analysis of Subliminal Learning reveals how student models acquire task capabilities from auxiliary teacher outputs via cross-task kernels.
datasette-explain 0.2.2 adds query plan visualization for read-only stored-query pages.
CTRL decouples LLM semantic reasoning from quantitative prediction in time series forecasting via frozen backbone and agent-based residual control.
Formal analysis of exact cost functions for self-calibrating monitors under online threshold adaptation.
World models trained on one action parameterization (absolute vs. delta) suffer 2.6–13.4× retrieval degradation when inference uses different parameterization.
End-to-end RL method using proximal residual value functions for two-timescale planning and real-time resource allocation.
Formal methods for fact-checking produce warrants (justification records) required by Digital Services Act and AI Act, beyond binary verdicts.
Diagnostic audit of Bayesian graph alignment convergence on 240+ exact and larger graph pairs reveals marginal disagreement gaps.
ChemCLIR-Bench: multilingual cross-lingual IR benchmark for chemical patents across five languages from Google Patents and EPO.
Temporal hierarchy forecasting reconciles hourly electricity prices and intraday spreads, improving accuracy up to 19.7% for battery arbitrage.
Mixture-learning framework recovers latent confounders as heterogeneity sources in observational data for causal estimation.
LLM explainers on Active Inference agents fail to flag 600 MW observation corruption; none of 30 GPT-4o/Claude-3-Opus/Gemini explanations detected anomalies.
Bayesian reformulation of ordinal regression with preference elicitation via card-insertion interface for decision modeling.
Meta's Muse is apparently an effective AI assistant, but one that's a little creepy. Part of that is because of its new Mac app, which can access Messages, Calendar, and Notes. But for all its smarts, Muse doesn't actually know how to describe itself. Jason Aten, a contributing editor at Inc Magazine, posted on Threads screenshots of an interaction he had with Muse in which the assistant asks him some questions about a conversation he was having in Messages. The problem is that Aten says he didn't give Muse access to his messages. When asked how it knew about the contents of his messages, Mus...
Without buyouts, Flock would "almost certainly" need to lay off staff.
Neural operator network (SDC-GON) for PDE solving using Green's functions with singular decomposition and consistency regularization.
Euston: 8B model trained to reject false mathematical claims instead of proving them, addressing reasoning LLM sycophancy.
Critique of general LLM rankings: benchmark saturation, data contamination, commercial bias, and task-specific evaluation gaps.
Trump claimed, without evidence, that the AI backlash is a Democratic hoax.
datasette-auth-github reaches 1.0 with session persistence fix for mobile browsers.
HGNN-based cross-modal knowledge transfer for low-resource speech representation learning in Yemba language.
CROSS-MAP: privacy-preserving LLM inference via semantic decoupling while preserving structural reasoning cues.
HGNN approach for low-resource speech-text multimodal alignment without large training datasets.
U-Net post-processing pipeline for scientific data compression, correcting pixel-space residuals in RVQ-based systems.
K-TRAIL: diffusion-guided EM/RF circuit layout generation using derivative-free Kalman filtering with black-box simulators.
GrapeSplat: feed-forward 3D Gaussian splatting from unposed images using voxel-aligned geometry without per-scene optimization.
Chronologic benchmark measures LM accuracy on historical English contexts 1831-1930 via pairwise comparison against multiple ground truths.
Conformal calibration framework converts point predictors into distribution-free uncertainty representations for robust downstream decision-making.