Autoregressive Boltzmann Generators
Autoregressive Boltzmann Generators improve molecular sampling efficiency beyond normalizing flows by removing invertibility constraints.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Autoregressive Boltzmann Generators improve molecular sampling efficiency beyond normalizing flows by removing invertibility constraints.
Study quantifies alignment between sequence probability and correctness across LLM decoding methods, models, and benchmarks.
Error-Conditioned Neural Solvers use learned correction instead of gradient descent to improve PDE solutions and extrapolation.
Modular open-source NLP pipeline extracts political-elite networks from multilingual text using joint entity-relation extraction.
Study analyzes BEACON method for low-resource entity matching under varying data constraints and domain awareness.
Language-based digital twin framework using LLMs detects Mild Cognitive Impairment via stylometric analysis and conversation patterns.
PEEU method trains open-source MLLMs for GUI task planning via autonomous environment exploration and hindsight experience replay.
MMBench2 dataset and study show world model hallucinations concentrate in low-coverage state-action regions and are preventable.
Despite ChatGPT's commanding market lead, consumers who pay for AI have been increasingly choosing Anthropic's Claude, data shows.
Sparse autoencoders with sparsity regularizers improve interpretability of vision model activations beyond top-k architectural constraints.
German Central Bank applies LLM-based information extraction to verify securities collateral eligibility from prospectuses, replacing rigid NER.
Gradient equilibrium in online optimization is algorithmically equivalent to Blackwell approachability, clarifying its position in online learning.
Taxonomy of indirect linguistic encoding (algospeak, euphemisms) mechanisms enables LLM detection of camouflaged social media expressions.
Context-aware translation cascades improve multilingual LLM reasoning by retaining original query and reasoning traces across translation stages.
Transfer learning framework combines lightweight physics simulations and autoencoders for structural damage diagnosis with limited experimental data.
Analysis of 15,000 reviews of AI healthcare chatbots identifies reliability, UX, and billing breakdowns; privacy concerns linked to lowest satisfaction.
datasette-export-database 0.3a2 released with dependency constraint fix for Datasette 1.0a27 compatibility.
Fast algorithm learns truncated high-dimensional Gaussians with optimal O(d²/ε²) sample complexity, improving on prior FOCS'24 results.
Analog Interaction Systems framework enables generative modeling on analog hardware (coupled oscillators, Ising machines) with physics-constrained dynamics.
RLAIF framework generates portable job search queries abstracted from seeker identifiers; addresses reward hacking in LLM-as-judge optimization.
Study identifies co-failure ceiling limiting accuracy gains in multi-model LLM routing/voting systems across 67 frontier models; proposes beta metric beyond pairwise correlation.
Empirical study of prompt injection attacks on LLM-based résumé screening; shows effectiveness collapses as injection adoption increases.
Compares simulation-based inference vs MCMC for Bayesian calibration of epidemiological COVID-19 models; domain-specific to infectious disease forecasting.
Theoretical work on identifiability and sample complexity bounds for learning ODEs from solution data; scientific ML focus, limited to core AI architecture audience.
Challenges scaling assumption in time-series forecasting; shows Ridge regression with optimized preprocessing rivals large transformers at lower cost.
Proposes physically-informed world model for Earth observation satellite forecasting under weather conditioning; domain-specific application.
Diagnostic framework decomposing LLM difficulty with historical text into tokenization cost, surprisal, and robustness across Italian datasets.
Introduces annotated dataset of manipulative betting ads from Instagram/Reddit for detection; dataset contribution limited to niche domain.
Ribbon: scalable uncertainty quantification method replacing bootstrap resampling with influence-function linearization for high-dimensional models.
E-TTS: embodied test-time scaling framework for robotic manipulation incorporating reasoning and historical context over long-horizon tasks.