7 papers
When Can Test-Time Adaptation Help Zero-Shot CT Vision-Language Models?
Ailar Mahdizadeh, Puria Azadi Moghadam, Xiangteng He +1
3D CT vision-language models (VLMs) classify abnormalities from text prompts in a zero-shot manner, enabling cross-institution deployment where labels are scarce and clinical tasks…
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
Ailar Mahdizadeh, Puria Azadi, Muchen Li +2
Streaming video understanding with large vision-language models (VLMs) requires a compact memory that can support future reasoning over an ever-growing visual history. A common sol…
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
Kun Yuan, Min Woo Sun, Zhen Chen +7
There is a growing interest in developing strong biomedical vision-language models. A popular approach to achieve robust representations is to use web-scale scientific data. Howeve…
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
Xiangteng He, Shunsuke Sakai, Shivam Chandhok +5
Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In p…
InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
Shunsuke Sakai, Xiangteng He, Chunzhi Gu +2
Despite the remarkable success, recent reconstruction-based anomaly detection (AD) methods via diffusion modeling still involve fine-grained noise-strength tuning and computational…
SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
Ailar Mahdizadeh, Puria Azadi Moghadam, Xiangteng He +3
Vision-language models (VLMs) have demonstrated strong cross-modal capabilities, yet most work remains limited to 2D data and assumes binary supervision (i.e., positive vs. negativ…