2 citations · 2 across the 1 of their papers we have counts for
10 papers · 1 filter
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition
Yuhuan Yang, Chaofan Ma, Zhenjie Mao +3
Video understanding is a complex challenge that requires effective modeling of spatial-temporal dynamics. With the success of image foundation models (IFMs) in image understanding,…
LoRKD: Low-Rank Knowledge Decomposition for Medical Foundation Models
Haolin Li, Yuhang Zhou, Ziheng Zhao +5
The widespread adoption of large-scale pre-training techniques has significantly advanced the development of medical foundation models, enabling them to serve as versatile tools ac…
Knowledge-enhanced Visual-Language Pretraining for Computational Pathology
Xiao Zhou, Xiaoman Zhang, Chaoyi Wu +3
In this paper, we consider the problem of visual representation learning for computational pathology, by exploiting large-scale image-text pairs gathered from public resources, alo…
ReMamber: Referring Image Segmentation with Mamba Twister
Yuhuan Yang, Chaofan Ma, Jiangchao Yao +3
Referring Image Segmentation~(RIS) leveraging transformers has achieved great success on the interpretation of complex visual-language tasks. However, the quadratic computation cos…
Reprogramming Distillation for Medical Foundation Models
Yuhang Zhou, Siyuan Du, Haolin Li +3
Medical foundation models pre-trained on large-scale datasets have demonstrated powerful versatile capabilities for various tasks. However, due to the gap between pre-training task…
Large-scale Long-tailed Disease Diagnosis on Radiology Images
Qiaoyu Zheng, Weike Zhao, Chaoyi Wu +7
Developing a generalist radiology diagnosis system can greatly enhance clinical diagnostics. In this paper, we introduce RadDiag, a foundational model supporting 2D and 3D inputs a…