PipeNetwork/minimax-h3-mlx
MiniMax releases H3, a multimodal generative model generating 15-second video from text/image/audio input; MLX port enables inference on Apple Silicon.
Every story tagged with this topic, ordered by date.
MiniMax releases H3, a multimodal generative model generating 15-second video from text/image/audio input; MLX port enables inference on Apple Silicon.
ParVL framework scales multimodal LLMs via parallel computation reuse between ViT and LLM, enabling task-specific optimization without fixed component allocation.
Agogic demonstrates music tokenization—not model scale—drives text-to-symbolic-music generation quality, using Qwen backbone with seven representations and Frechet Music Distance.
Video-DeepResearch extends multimodal agents to video streams; identifies modality bias and knowledge leakage bottlenecks in current models.
EcoFrame enables training-free adaptive frame scheduling for long video understanding in VLMs, using inference feedback to adjust sampling budgets.
CARE-X unifies chest X-ray report generation with classification, localization, and measurement for clinical VLM utility.
BanglaWild benchmark evaluates 15 VLMs and OCR systems on 2,535 in-the-wild Bengali scene text images.
CheMatE: ModernBERT-based embedding model for joint SMILES-NLP representation learning in chemistry, addressing domain overfitting.
Geo-Embed: unified multimodal embedding model for urban/geospatial tasks spanning street-view imagery, remote sensing, text, and temporal change.
Narrative review of AI-based sound effect generation across input modalities: text, visual, audio, and multimodal for digital applications.
Self-augmentation method for MLLMs using model failure signals to generate targeted image augmentations without external supervision.
Pattern Completion Bias benchmark measuring how repeated UI patterns degrade multimodal LLM accuracy on screenshot-to-code fill-in-the-blank tasks.
RAPO framework addresses catastrophic forgetting in continual multimodal LLM post-training via explicit dual-channel risk governance.
Study proposes explicit modality reliability modeling for incomplete multimodal sentiment analysis across text, audio, vision.
UEmbed is a decoder-only multimodal model producing sparse lexical and dense embeddings in one pass, extending learned sparse retrieval to multimodal RAG without auxiliary modules.
Study shows Qwen-VL achieves 87.3% accuracy on vehicle damage classification but fails spatial grounding for fine-grained defects; proposes dedicated segmentation layer.
Pinterest deploys vision-language models for automated relevance evaluation in search, reducing human annotation cost at scale.
Benchmark of 11 audio-language models and classifiers on closed-set sound source identification (2,242 clips, 23 classes).
MCR-GRPO assigns marginal box contributions in multimodal RL for structured visual perception tasks without response-level broadcast.
Attribute-level unlearning for MLLMs enables fine-grained removal of sensitive identity information while preserving model utility.
PMMC compiles multimodal memory at consolidation time for LVLM agents to preserve image-text binding and temporal updates without query-time overhead.
FriendBench evaluates 26 multimodal LLMs on dyadic familiarity inference from video; top models match human performance but via different mechanisms.
Vision-language models exhibit sycophancy that undermines epistemic vigilance in cooperative tasks; detects inconsistencies with prior beliefs.
MolGVR framework adds verification and refinement to text-to-molecule generation to enforce chemical constraints and correct structural violations.
QR-STT method crafts thermal adversarial patterns to steer infrared vision-language models via optimized thermal states; demonstrates robustness risk.
Analytic memory abstraction for multimodal agents enabling filtering, aggregation, and temporal reasoning over accumulated observations.
Autoregressive speech generation using low-frame-rate high-dimensional continuous tokens to balance stability and reconstruction fidelity.
ReToken learnable embedding improves vision-language models on visual retrieval by selecting sparse tokens from KV cache.
VAD counterfactual algorithm isolates visually attributable components in multimodal on-policy distillation corrections.
DualG-MRAG decouples macro reasoning and micro-matching in multimodal retrieval-augmented generation to handle complex multi-hop tasks.
ScaFE uses LLM-generated feature programs to classify pathological scars (keloids vs. hypertrophic) from clinical photos with limited labeled data and local data governance.
EndoCLIP, a vision-language foundation model trained on 280k colonoscopy reports, improves lesion detection and report generation.
Perception-Correction Distillation isolates perception failures in multimodal reasoners using teacher-student disagreement without labeled ground truth.
Google DeepMind releases Gemini Robotics ER 2 with video understanding, task orchestration, and multi-robot coordination capabilities.
PathView-Bench evaluates multimodal LLMs on fine-grained, multiscale pathology image understanding with spatial annotations across 23 datasets.
ObjectStream organizes streaming video context around persistent latent objects as memory anchors instead of token or temporal importance.
MonoVoc decouples geometry and semantics for lightweight monocular open-vocabulary 3D Gaussian scene understanding via training-free pipeline.
Theia auto-captions and validates Incidents1M, a multimodal disaster dataset for vision-language models via data-free knowledge distillation.
EgoGenesis synthesizes egocentric manipulation videos using geometry-aware conditioning and projective memory to expand embodied AI training data.
Black-box evaluation reveals commercial multimodal content moderation APIs vulnerable to simple image transformations.
RRM augments multimodal memory graphs with reflection mechanisms for long-horizon video reasoning agent adaptation.
OPLD uses on-policy latent distillation to train flexible latent visual reasoning without external trace supervision.
MMAC: 5,638-clip benchmark for audio captioning across 15 evaluation dimensions, targeting fine-grained free-form descriptions for AudioLLMs.
SciFigQual-Bench: benchmark for scientific figure quality assessment using full-manuscript context; evaluates caption alignment and visual misleadingness.
Single-beat cuffless blood pressure estimation via ear-PPG and ECG with lightweight hybrid learning for wearables.
Visual Credit Audit (VCA) measures whether multimodal models genuinely use image information vs. text-only reasoning in spatial benchmarks.
SciFigAlign evaluates scientific figures by alignment of visuals with manuscript claims, addressing peer-review assessment beyond generic image quality.
Progressive Multimodal Alignment mitigates projector drift in continual instruction tuning of vision-language models.
BioVLN simulation platform for visual-language navigation in biomedical labs with instrument-specific approach constraints and safe clearance requirements.
DuPLeR dual-path LLM reasoning framework combines multimodal and LLM priors for few-shot knowledge graph completion while filtering hallucinated evidence.