Score: 1

EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography

Published: December 13, 2025 | arXiv ID: 2512.12107v1

By: Yuheng Li , Yue Zhang , Abdoul Aziz Amadou and more

Potential Business Impact:

Helps doctors understand heart scans faster.

Business Areas:

Image Recognition Data and Analytics, Software

Echocardiography is the most widely used imaging modality in cardiology, yet its interpretation remains labor-intensive and inherently multimodal, requiring view recognition, quantitative measurements, qualitative assessments, and guideline-based reasoning. While recent vision-language models (VLMs) have achieved broad success in natural images and certain medical domains, their potential in echocardiography has been limited by the lack of large-scale, clinically grounded image-text datasets and the absence of measurement-based reasoning central to echo interpretation. We introduce EchoGround-MIMIC, the first measurement-grounded multimodal echocardiography dataset, comprising 19,065 image-text pairs from 1,572 patients with standardized views, structured measurements, measurement-grounded captions, and guideline-derived disease labels. Building on this resource, we propose EchoVLM, a vision-language model that incorporates two novel pretraining objectives: (i) a view-informed contrastive loss that encodes the view-dependent structure of echocardiographic imaging, and (ii) a negation-aware contrastive loss that distinguishes clinically critical negative from positive findings. Across five types of clinical applications with 36 tasks spanning multimodal disease classification, image-text retrieval, view classification, chamber segmentation, and landmark detection, EchoVLM achieves state-of-the-art performance (86.5% AUC in zero-shot disease classification and 95.1% accuracy in view classification). We demonstrate that clinically grounded multimodal pretraining yields transferable visual representations and establish EchoVLM as a foundation model for end-to-end echocardiography interpretation. We will release EchoGround-MIMIC and the data curation code, enabling reproducibility and further research in multimodal echocardiography interpretation.

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence

CV and Pattern Recognition

Helps doctors find cancer faster with sound pictures.

18 Sep 2025 1

91%

Video CLIP Model for Multi-View Echocardiography Interpretation

CV and Pattern Recognition

Helps doctors understand heart videos better.

26 Apr 2025 0

89%

EchoAgent: Guideline-Centric Reasoning Agent for Echocardiography Measurement and Interpretation

CV and Pattern Recognition

Helps doctors understand heart videos better.

17 Nov 2025 1

View PDF Login to Bookmark

Page Count

27 pages

EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography

Helps doctors understand heart scans faster.

Technical Abstract

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence

Video CLIP Model for Multi-View Echocardiography Interpretation

EchoAgent: Guideline-Centric Reasoning Agent for Echocardiography Measurement and Interpretation