4 papers · 1 filter
StackTok: Accelerating VLMs Inference with Budget-Adaptive Visual Token Selection
Zhenbin Wang, Lei Zhang, Lituan Wang +3
Increasing image resolution produces ever-longer visual-token sequences in vision-language models (VLMs), substantially raising their inference cost. To reduce this overhead withou…
SonoCLIP: Mask-Guided Region-Aware Vision-Language Pretraining for Fetal Ultrasound Analysis
Hang Su, Chao Sun, Zhaofan Li +3
Vision-language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging have recently emerged, the domain remains…
Physics-Driven 3D Gaussian Rendering for Zero-Shot MRI Super-Resolution
Shuting Liu, Lei Zhang, Wei Huang +2
High-resolution Magnetic Resonance Imaging (MRI) is vital for clinical diagnosis but limited by long acquisition times and motion artifacts. Super-resolution (SR) reconstructs low-…
DualFete: Revisiting Teacher-Student Interactions from a Feedback Perspective for Semi-supervised Medical Image Segmentation
Le Yi, Wei Huang, Lei Zhang +3
The teacher-student paradigm has emerged as a canonical framework in semi-supervised learning. When applied to medical image segmentation, the paradigm faces challenges due to inhe…