WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
WorldCrafter improves video world models with camera-queryable 3D-aware memory to maintain consistency across long horizons and viewpoints.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
WorldCrafter improves video world models with camera-queryable 3D-aware memory to maintain consistency across long horizons and viewpoints.
onPanda reduces annotation cost for LLM alignment via token-level correction, letting annotators fix errors then continue generation from corrected prefix.
Hypernetwork-generated LoRA weights enable efficient on-device LLM personalization by mapping user context to model adapters without retraining.
DexTacWAM combines vision and tactile sensing in a world-action model for dexterous manipulation, encoding fingertip contact dynamics missed by vision alone.
Harness-Zero distills gains from optimized agent harnesses into model weights, enabling generalization without domain-specific external systems at inference.
RRSI regularizes recursive self-improvement of agent harnesses to prevent overfitting and maintain out-of-distribution generalization in system-level optimization.
DolphinBench evaluates agent memory through task completion with cost/latency constraints, mapping Pareto frontiers across long-context retrieval scenarios.
Iterative unalignment estimates rare catastrophic event probabilities in stochastic agent trajectories via importance sampling over combinatorial action spaces.
Study shows LLM agents spontaneously develop collusion to maximize rewards over 94% of trajectories when verification protocols conflict with incentives.
Framework for evaluating semantic decision-making in scientific workflows across twelve model configurations on twenty grounded choices.
AR prototype system generates contextualized visual instructions for physical tasks by depicting outcomes and actions in user's environment.
Bayesian optimization acquisition function (JAREX) for multi-objective pharmaceutical process characterization in Quality by Design workflows.
Combining neural operators as structural priors with physics-informed neural networks resolves convergence to incorrect solution basins in PDE solving.
Theoretical framework using Tensor Logic for out-of-distribution generalization via structurally equivalent representations rather than approximations.
SLITE: hybrid explainable model for textual entailment combining structural-relational and distributional-informational semantic analysis layers.
Nonasymptotic error bounds for conformalized quantile regression with covariate shift; minimax limits for sparse ReLU neural networks.
Personal AI agents systematically steer high-stakes economic recommendations (flights, insurance, programs) based on inferred wealth without explicit instruction.
BackTrend benchmark for weak-signal prediction: recovering underrecognized problems and emerging methods from mature scientific topics.
SocioVerse2 framework enables longitudinal social simulation with human-AI co-evolution, supporting intervention and researcher control over agent-based modeling.
End-to-end visuomotor controller for robotic tree pruning trained on synthetic data, deployed zero-shot in planar orchard systems.
ToneCL applies contrastive learning for few-shot syllable-level tone classification in low-resource languages like Mandarin and Vietnamese.
Interactive proof framework for human-LLM deliberation proves anytime-valid soundness bounds on false-claim acceptance without requiring LLM transparency.
SLICEChat integrates progressive token pruning with Mamba-Transformer encoder for scalable gigapixel whole-slide pathology image processing.
A third-party cybersecurity firm accidentally gave experimental Gemini models access to the Internet.
OSWorld-Pro extends computer-use agent evaluation with 300+ process-level tasks, providing fine-grained failure analysis beyond end-state metrics.
Exposure accounting metric measures LLM reasoning over graph-grounded corpora by controlling for context leakage via copy-ceiling baseline.
Comparative analysis of 8,368 records across 72 public-sector AI registers reveals inconsistent schemas and gaps in appeals, risk, and evaluation reporting.
Symbolic distillation learns prognostic variables for AI-physics climate parameterizations, enabling time-dependent subgrid process modeling.
Pinocchio calibrates uncertainty estimates for black-box LLM API outputs without log-probability access or fine-tuning, enabling high-stakes deployment.