Dynamic Resource Allocation for Ensemble Determinization MCTS
Dynamic Resource Allocation for Ensemble Determinization MCTS improves search efficiency in high-uncertainty adversarial games via adaptive determinization.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Dynamic Resource Allocation for Ensemble Determinization MCTS improves search efficiency in high-uncertainty adversarial games via adaptive determinization.
Spectral predictability indices conflate series complexity with context value; phase structure, not spectrum alone, determines benefit of retrieval/pretraining.
Watermark forensics theory shows attribution, payload extraction, and edit-localization form a sample-complexity ladder quantified by information profiles.
DeepMind CEO Demis Hassabis is proposing an AI "standards body" modeled after FINRA, to test frontier models and develop best practices for their release.
LLM plan evaluators exploit deletion non-monotonicity—rewarding explicit steps' removal via score-seeking optimization on venture routing tasks.
Method to certify LLM honesty under incentive pressure via counterfactual report mediators that resist manipulation while remaining responsive to evidence.
Neural-symbolic framework for automatic generation of multimodal analytic geometry problems, addressing precision gaps in diagram generation for math reasoning.
A group of 26 former Meta employees is suing the company over claims that it used AI tools to unfairly target workers on leave with layoffs, as reported earlier by Reuters. In the lawsuit, the employees allege Meta determined which workers to dismiss based on performance data collected by a "constellation" of internal AI tools, but failed to exclude those on parental or medical leave from its ranking system: The result was that employees who took protected leaves were disproportionately selected for layoff, based on scoring that not only failed to account for their protected leaves, but in ef...
Ensemble Controlled-flow Filter for data assimilation in implicit observation systems using energy tilt and stochastic flows.
LLM predictions flip substantially under task-irrelevant context despite stable aggregate accuracy, exposing fragility masked by benchmark metrics.
Placebo-controlled methodology (PoPE) to measure whether small code LLMs can actually use execution error feedback to repair code.
Engineering use of AI forecasting models requires not only high nominal accuracy but also predictable behavior under uncertain inputs. In photovoltaic (PV) forecasting, this requirement is especially challenging because numerical weather prediction (NWP) errors are temporally correlated, state dependent, and physically coupled across variables. Existing evaluations, however, often rely on perfect forecast assumptions or simplistic perturbations that do not reflect these characteristics. This study presents a physically constrained robustness evaluation framework based on simulation, using vir...
Datasette 1.0a37 alpha release with permissions system improvements and plugin compatibility fixes.
Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interaction resource. Building such a resource is challenging for three reasons: cross-platform hotel names must be reconciled before interactions are comparable; quality must be audited with reproducible metrics rather than ad hoc cleaning; and public release must preserve privacy while remaining benchmarkable under realistic cold-start conditions. We introduce ViHoRec, a quality-controlled Vietnamese hotel recommendation dataset of 18{,}267 interactions between 6{,}832 users and 560 ...
The new Google image search will use your "unique interests" to create an always-updated gallery.
Instagram head Adam Mosseri believes companies will eventually need to manage AI token spending the same way they manage payroll or other operating expenses, predicting that engineers could soon face limits on how much they spend using AI tools.
We study the online binary sequential calibration problem. A recent breakthrough by \citet{dagan2024breaking} overcomes the classical \(T^{2/3}\) barrier for calibration error. Building on this result, we present an efficient randomized forecaster that achieves an expected calibration error \(O(T^{2/3-\varepsilon})\) for some constant \(\varepsilon>0\). Our forecaster combines the \textsc{SPR-Calibration} procedure \citep{dagan2024breaking} with an outer Blackwell-style correction layer. The \textsc{SPR-Calibration} procedure controls calibration with respect to a surrogate sequence of condit...
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve... Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve build issues, launch experiments, monitor execution, analyze metrics, and summarize results. For reinforcement learning (RL) research, this matters because meaningful metrics often appear only after the essential experiment infrastructure… Source
Google Images celebrates 25 years with retrospective on visual search milestones and new content exploration features.
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning... What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning models to production video tasks, developers often lose days to data formatting, container setup, training scripts, baseline evaluation, and hyperparameter sweeps before they even know whether post-training improves accuracy. Source
Now, when users navigate to Google Images, they'll see a "For You" gallery of images tailored to their interests and browsing history.
Google is announcing a big change to the Google Images homepage in honor of the platform's 25th anniversary this week. Instead of a mostly blank page with a search bar, the homepage will soon show you a bunch of images that it thinks you might like before you even start searching. The company says the new "browseable" homepage features a "dynamic, immersive gallery of images from across the web - updated in real time and intelligently tailored to your unique interests." Based on images Google has shared, the layout reminds me of platforms like Pinterest and Imgur that stuff a lot of images in...
In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes their training sample inefficient. Though in most PAMDP environments explicit but incomplete knowledge (e.g., rules, safety constraints, or expert heuristics) is available, it is rarely directly used to increase the sample-efficiency of training Reinforcement Learning agents. We step into this gap...
Stochastic-process models are, as a rule, far easier to simulate than to condition. Non-linear observations, non-Gaussian likelihoods, black-box information, and global constraints all induce intractable conditional laws, requiring bespoke, model-specific constructions. We introduce LatentFlow, a single framework for conditioning stochastic processes, with no learned neural approximations and no training. Our starting point is to write the stochastic process as the deterministic image of a tractable latent innovation, $f_0 = T_{\vartheta}(ξ_0)$, with $ξ_0$ sampled from a simple reference dist...
CoCo loss function for learning normalized embeddings with intra-class collapse and inter-class contrast, with theoretical analysis vs. cross-entropy.
Physics-informed fall detection framework using dual-LTC architecture for CoM and BoS subsystems on edge platforms.
Theoretical analysis showing Randomized Hamiltonian Monte Carlo achieves accelerated mixing time for log-concave distributions.
MemOps benchmark for granular evaluation of LLM agent memory operations in long-horizon conversations beyond final QA correctness.
UR-VC method for unsupervised correction of time-derived progress labels in robot learning with contact-rich manipulation.
Energy-based learning framework for inverse problems in tensegrity structure form-finding and physical property prediction.