7 papers · 1 filter
Open-Ended CT Volume Segmentation with Weak Supervision from Language
Sanjay Subramanian, Junwei Yu, Zirui Wang +5
We introduce a method for training a text-conditioned segmentation model for CT scans, which combines voxel-level supervision with coarse but scalable slice-level supervision from…
Stateful Visual Encoders for Vision-Language Models
Zirui Wang, Junwei Yu, Adam Yala +3
Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes. However, in existing open-weight VLMs, vis…
Pillar-0: A New Frontier for Radiology Foundation Models
Kumar Krishna Agrawal, Longchao Liu, Long Lian +11
Radiology plays an integral role in modern medicine, yet rising imaging volumes have far outpaced workforce growth. Foundation models offer a path toward assisting with the full sp…
Describe Anything: Detailed Localized Image and Video Captioning
Long Lian, Yifan Ding, Yunhao Ge +8
Generating detailed and accurate descriptions for specific regions in images and videos remains a fundamental challenge for vision-language models. We introduce the Describe Anythi…
Rethinking Patch Dependence for Masked Autoencoders
Letian Fu, Long Lian, Renhao Wang +6
In this work, we examine the impact of inter-patch dependencies in the decoder of masked autoencoders (MAE) on representation learning. We decompose the decoding mechanism for mask…
TULIP: Towards Unified Language-Image Pretraining
Zineng Tang, Long Lian, Seun Eisape +6
Despite the recent success of image-text contrastive models like CLIP and SigLIP, these models often struggle with vision-centric tasks that demand high-fidelity image understandin…