Parallel Decoding Distillation for Fast Image and Video Generation
Parallel Decoding Distillation accelerates video/image diffusion models via trajectory-based training, avoiding mode collapse and motion loss of score distillation approaches.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Parallel Decoding Distillation accelerates video/image diffusion models via trajectory-based training, avoiding mode collapse and motion loss of score distillation approaches.
Analysis of Sharpness-Aware Minimization and Muon optimizer investigates matrix-aware geometry for robust generalization under parameter perturbations.
Empirical evaluation of nine Tabular Foundation Models reveals robustness gaps under distribution shift versus ensemble tree methods across diverse pre-training strategies.
Study shows LLM-generated Kubernetes security patches improve when prompted with runtime service topology context rather than isolated KSPM findings.
MemLens provides value-aware memory management for LLM agents with interactive analytics, treating memory records as first-class objects to reduce redundancy.
Framework formalizes co-drift failure prediction in self-driving networks by treating monitoring, analytics, and actuation as coupled macro-intents.
GARI introduces representation-level soft-equivariance design for generic sequence backbones using aligned generator views across modalities.
Deep RL for quadcopter control via physics-aware end-to-end learning with high-fidelity actuator dynamics simulation.
GARFIELD: probabilistic latent model for predicting distributions over future scene kinematics from partial observations.
OpenAI field report documents how AI coding agents accelerate scientific computing workflows in genomics and adjacent domains.
RL for code optimization via learnable execution time rewards and DMC-Optim framework to overcome measurement noise and sparsity.
Quasi-SVD: GPU-parallelized differentiable matrix factorization with Lie-constrained orthogonality for real-time medical imaging.
PRISM-AH: multimodal framework for detecting ambivalence/hesitancy in health behavior video via cross-modal conflict reasoning.
Kontra: taxonomy and detection framework for cross-modal knowledge inconsistencies in Wikipedia, Wikidata text/tables/KGs.
Solver-guided LLM framework for instance-wise operations research formulation selection in multi-warehouse inventory allocation.
Polistemics: theory-grounded benchmark for evaluating LLMs as political information mediators via epistemic modesty standard.
MODUS: decoder-only any-to-any multimodal model supporting arbitrary modality combinations with pre-trained model priors.
ClinPRISM: cost-effective multimodal LLM for QA over irregular clinical time series via sparsity-aware encoding.
ClinMM-Bench evaluates multimodal LLMs on multi-turn clinical diagnosis with progressive information disclosure, mirroring real-world diagnostic reasoning.
Empirical evaluation of flow matching, DDPM, score-SDE, and VAE on non-stationary Gaussian random fields to assess whether DGMs capture underlying spatial processes.
Survey of face de-identification techniques spanning post-capture processing and capture-time privacy preservation for computer vision tasks.
Extension of dtControl2 decision tree tool with epsilon-bounded optimality tradeoff for explainable MDP policy representation in large state spaces.
Benchmark of six VLMs (Gemini, GPT-4V, Qwen, Gemma, Llama, Ministral) on zero-shot anomaly detection for game geometry clipping in agent-driven QA.
Penelope framework localizes latent recurrent computation to decoder intervals for efficient structured reasoning without extending autoregressive chain-of-thought output.
AgentToolMO proposes 3GPP NRM information model for cross-vendor AI agent tool trust management with graduated enforcement and cascade propagation bounds.
Vision-Language-Action model framework using SAM3D as frozen 3D teacher to align fine-grained object representations for robotic manipulation under occlusion.
AnnoBench introduces evaluation framework for visualization annotation generation systems balancing visual, semantic, and stylistic constraint satisfaction.
Input-only prompt optimization suppresses evaluation-awareness latents in LLMs via GCG-style token search, threatening safety evaluation validity if undetected.
Interactive Reward Agent uses environment-state verification to evaluate GUI task completion, enabling reliable reward signals for agentic scaling.