Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
Evidence-Grounded Social Persona Panel evaluates generative UI quality across psychologically diverse user representations.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Evidence-Grounded Social Persona Panel evaluates generative UI quality across psychologically diverse user representations.
Short-term graph memory module optimizes molecular search by pre-screening candidates under fixed oracle budget constraints.
Analysis of LLM hidden states using aggregator/differentiator metrics to measure how tokens consolidate text representation and metaphorical transport across positions.
UNICON foundation model demonstrates numerical intelligence via in-context learning across scientific/social datasets, generalizing beyond language-only reasoning.
The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and other non-human identities across cloud environments.
Physics-inspired KSSE method replaces dense CNN classifiers using sparse-graph spectral embedding and Ising models for image classification.
READII-2-ROQC framework uses negative controls to detect volume confounding in radiomics and imaging foundation model biomarkers.
QAdapt neural pre-decoding framework adapts to nonstationary hardware noise in quantum error correction via noise-adaptive decoders.
Study of derived-feature over-trust (DFOT) in LLMs using physiological sensing; proposes privileged-modality reliability evidence to mitigate misapplication.
WIDE framework enables token-level dynamic width pruning for adaptive LLM inference, improving efficiency while preserving accuracy under aggressive sparsity.
QQWorld replaces Epps-Pulley objective with quantile-quantile matching to regularize latent world model distributions for improved planning.
Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We... Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We routinely see 8% to 12% gaps between partner deployments and the corresponding NVIDIA reference architecture (RA) on the same workload, same model, same global batch size. The cause is often a stack of configuration choices in the kernel… Source
Query complexity analysis of windowed thinning for bouncy particle and Zigzag samplers under strongly convex/smooth potentials.
PACE hierarchical framework applies LLMs to parent-order execution in algorithmic trading, generalizing across market conditions without task-specific training.
LLM Chat Completions Server 0.1a0 release adds OpenAI-compatible API endpoint for local model serving.
Meta says AI is making it dramatically easier to build and launch new consumer apps, with CEO Mark Zuckerberg telling investors the company has more new consumer products on the way following a recent wave of releases for Facebook Groups, Marketplace sellers, Instagram, and gaming.
LLM 0.32rc1 introduces content-addressable message storage and schema redesign enabling conversation branching.
British AI neocloud Nscale is buying software startup Anyscale, which helps companies scale their AI workloads across data centers and servers.
Federated learning method using encrypted metadata clustering to balance privacy, communication, and computation without sacrificing one dimension.
Perception-Correction Distillation isolates perception failures in multimodal reasoners using teacher-student disagreement without labeled ground truth.
Google DeepMind releases Gemini Robotics ER 2 with video understanding, task orchestration, and multi-robot coordination capabilities.
Anthropic publishes findings from cybersecurity evaluations testing AI model capabilities against real-world attack scenarios and incident response.
A new study estimates only 2,000 U.S. engineers have the expertise to deliver meaningful AI ROI, as enterprises race to hire forward-deployed engineers to implement AI at scale.
CARP reputation-penalty mechanism prevents LLM agents from fabricating product listings using complaint signals without access to ground truth.
Gap Index metric evaluates dimensionality reduction scatterplot quality by measuring distortion in empty regions alongside traditional point-distance metrics.
Fairness Pruning localizes demographic bias in GLU-MLP layers via differential neuron activations, enabling targeted bias mitigation in LLMs.
PathView-Bench evaluates multimodal LLMs on fine-grained, multiscale pathology image understanding with spatial annotations across 23 datasets.
Plus, a new deprecation policy ensures features aren't removed suddenly.
Budget-constrained human audit allocation for N LLM agents identifies miscalibration threshold where confidence-ranking underperforms random selection.