FleXray: Universal Clinical X-ray Segmentation
FleXray is a generalist X-ray segmentation model spanning whole-body anatomy for clinical analysis, addressing 2D projection ambiguity.
Every story tagged with this topic, ordered by date.
FleXray is a generalist X-ray segmentation model spanning whole-body anatomy for clinical analysis, addressing 2D projection ambiguity.
MMAP pretraining framework handles missing multimodal data and incomplete records for longitudinal Alzheimer's disease progression prediction.
Foundation model embeddings from screening mammograms encode pre-diagnostic tissue changes detectable before cancer diagnosis, varying by pretraining domain.
Joint recognition and translation fine-tuning via group relative policy optimization (GRPO) for LLM-based speech translation with Qwen2.5-Omni-3B.
Feature-wise linear modulation (FiLM) integrates radiomics with RenalCLIP foundation model for CT-based renal cell carcinoma classification.
CHiME-9 ECHI challenge evaluates multi-channel speech enhancement from noisy cafeteria conversations using Meta Aria and hearing aid microphones.
REVE detects audio hallucinations in audio-language models by reusing encoder states for efficient event verification without second forward pass.
VideoX-Qwen framework for instruction-driven video editing using paired data and adapted video-generation backbones.
Study identifies domain gap between simulated multi-speaker extraction training and real conversational speech with imbalanced speaker presence and enrollment mismatch.
AR prototype system generates contextualized visual instructions for physical tasks by depicting outcomes and actions in user's environment.
SLICEChat integrates progressive token pruning with Mamba-Transformer encoder for scalable gigapixel whole-slide pathology image processing.
Multimodal wearable sensing detects agitation in autistic youth via IMU, physiology, vocalization; medical/clinical application, not core AI research.
PrismGPT is a VLM framework for region-aware photo editing that produces structured plans without commercial tools.
Orthogonal knowledge distillation method for open-vocabulary audio-visual event localization that selectively weights teacher supervision by temporal-boundary reliability.
RA-MMEA framework assesses visual reliability and adaptively improves image representations for robust multi-modal entity alignment.
ChemCLIR-Bench: multilingual cross-lingual IR benchmark for chemical patents across five languages from Google Patents and EPO.
HGNN-based cross-modal knowledge transfer for low-resource speech representation learning in Yemba language.
HGNN approach for low-resource speech-text multimodal alignment without large training datasets.
QwenVLConnector medical VLM unifies classification, detection, counting, regression, and report generation via dense multi-layer Connector on Qwen2.5-VL.
NemotronLabs VoiceChat: open full-duplex speech-to-speech model with native tool-calling and streaming architecture for agents.
Benchmark framework evaluating explanation quality alongside accuracy for vision-language models in forensic face recognition tasks.
MIST multimodal model fuses genomic and histology data via attention for cancer survival prediction.
Study reveals vision-language models fail to meaningfully use patient ECGs in clinical prediction, proposing detection and mitigation strategies.
ReACT-TTS uses listener facial reactions to plan emotion and prosody in conversational speech synthesis.
Spoken Wikipedia Presentation Corpus extends ASR datasets with LLM-generated slides for multimodal speech recognition evaluation.
Samsone family of small audio language models (99M–356M params) achieving SOTA on edge inference benchmarks.
Configurable multi-stage vision pipeline for crop disease/pest diagnosis in Farmer.Chat, enabling threshold tuning and new disease/crop addition.
Paint-Anything enables hex-color control for image generation/editing using LM-based color semantics without specialized representations.
Video DeltaNet combines local Softmax and linear attention for efficient high-quality video generation by balancing fine-grained interactions with computational cost.
HerHealthEval benchmark: multilingual, register-sensitive evaluation of LLM understanding in women's health across English, French, Modern Standard Arabic.
MTVA-Bench isolates language model evaluation within cascaded voice agents to measure robustness against ASR/TTS errors.
Surface-based diffusion bridge synthesizes FDG-PET from MRI on cortical manifolds for dementia biomarker accessibility.
FedQoS enables asynchronous federated learning for multimodal sensors in smart vehicles with heterogeneous network constraints.
Astronex-World 1.0: open video world model with text/image conditioning, camera trajectories, and block-causal attention for real-time generation.
AVTrace benchmark suite with 34,114 examples diagnoses temporal reasoning in multimodal models across audio-visual synchronization and event localization.
VākQA introduces a 2,001-pair Telugu spoken QA benchmark with speech audio and validates LLM-based evaluation methods against human judgments.
PANORAMA addresses panoptic grounded captioning, combining dense captions with pixel-level segmentation for vision-language models.
MUSE benchmark evaluates vision-language models on multimodal artistic and cultural understanding for AI-assisted language education.
ReFigBench evaluates multimodal agents on reconstructing scientific figures as editable PowerPoint, isolating visual perception from planning and tool harness failures.
Causal analysis of VLM attention heads reveals general-purpose semantic features enabling OCR, with heads outputting interpretable tokens across diverse image regions.
Zero-shot cross-lingual handshape recognition transfers ASL phonological features to Catalan Sign Language via decomposed feature prediction.
SynAgent uses multimodal LLM agents to autonomously guide materials synthesis experiments with explicit reasoning and adaptive analysis.
VoiceTrace benchmark enables speaker-aware speech retrieval combining semantic content search with speaker identification from reference utterances.
Aligned Continuous Integrate-and-Fire framework compresses speech tokens efficiently for zero-shot SpeechLLM alignment without full-model fine-tuning.
ActiveScale framework enables vision-language-action models to perform active perception via coordinated model, data, and hardware designs for robotic manipulation.
Mixture-of-Bottleneck Experts reformulates video sentiment analysis as ordinal regression across text, audio, image modalities.
Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking speech-to-speech models with browser-based UI supporting voice interruption.
DELTA and TARQA improve table understanding via structured text over vision-language models; supports multilingual documents.
Google extends Gemma multilingual capabilities beyond text translation to preserve nuance across world languages.
Multi-judge ensemble approach for detecting hallucinated character spans in vision-language outputs; SHROOM-Visions shared task winner.