← Back to feed
2026-07-08multimodaldatavision

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen

PDF preview for MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
Read on arXiv →

Key claim

MedPMC improves medical multimodal model performance significantly.

In plain English

Imagine you're a clinician trying to make sense of a mountain of medical literature and images. You need to pull together information from various sources, but the data is often messy, incomplete, or just plain wrong. This is a common issue in medicine, where the quality of data can directly impact patient care. When researchers try to build models that understand this data, they often run into problems because the existing datasets lack the necessary fidelity and clinical validation. This is what's called data quality issues in multimodal models.

To tackle these challenges, the authors developed a system called MedPMC. This framework automatically curates high-quality image-text pairs from millions of articles in PubMed Central, which is a treasure trove of medical literature. By applying advanced techniques, MedPMC ensures that the curated data is not only vast but also relevant and reliable. For instance, it achieved impressive performance metrics in various tasks, such as detecting multi-panel figures and classifying medical images.

What sets MedPMC apart from previous efforts is its ability to significantly improve the performance of models trained on this data. For example, a model trained with MedPMC data outperformed existing models in medical visual question-answering tasks and retrieval tasks, showing that better data leads to better outcomes. This means that if you're building applications in healthcare that rely on understanding medical images and texts, using MedPMC could give you a substantial edge in accuracy and reliability.

Novelty
8.0/10

The introduction of MedPMC significantly enhances the quality of multimodal medical data curation.

Reliability
8.0/10

The paper provides strong empirical results across multiple benchmarks and a large dataset.

Deep reliability assessment

The methodology supports the creation of a high-fidelity dataset for medical multimodal models, but the claim of improving clinical settings may be overclaimed without extensive real-world validation.

Reproducibility

yes, the framework, corpus, benchmarks, and pretrained models are publicly released.

Key figure

Figure 2a shows the proportion of medical versus non-medical images in the MedPMC and BIOMEDICA datasets.

Benchmark results

~MedPMCF1: 93.2vs MedICaT keyword-matching+31.5%SOTA
~MedPMCF1: 96.5vs text-based approachesN/ASOTA
~MedPMCmAP: 89.8vs YOLO-OCR-D-0.4%
~MedPMCF1: 81.4vs PMC-OA pipeline+32.9%SOTA
~MedPMCF1: 96.5vs other vision modelsN/ASOTA