4 papers
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
Min Woo Sun, Alejandro Lozano, Javier Gamazo Tejero +8
Embedding vision-language models (VLMs) are typically pretrained with short text windows (<77 tokens), which forces the truncation of long-format captions. Yet, the distribution of…
Unified Supervision For Vision-Language Modeling in 3D Computed Tomography
Hao-Chih Lee, Zelong Liu, Hamza Ahmed +6
General-purpose vision-language models (VLMs) have emerged as promising tools in radiology, offering zero-shot capabilities that mitigate the need for large labeled datasets. Howev…
Text2CT: Towards 3D CT Volume Generation from Free-text Descriptions Using Diffusion Model
Pengfei Guo, Can Zhao, Dong Yang +9
Generating 3D CT volumes from descriptive free-text inputs presents a transformative opportunity in diagnostics and research. In this paper, we introduce Text2CT, a novel approach…
VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge
Vishwesh Nath, Wenqi Li, Dong Yang +22
Generalist vision language models (VLMs) have made significant strides in computer vision, but they fall short in specialized fields like healthcare, where expert knowledge is esse…