activity
20242026
collaborators

7 papers

cs.CV2026

SAMRI-3D: Adapting SAM2 for 3D MRI Segmentation with Global Volume Tokens

Zhao Wang, Wei Dai, Hongfu Sun +2

Foundation models such as Segment Anything Model 2 (SAM2) have transformed natural-image and video segmentation, and recent work has begun adapting them to medical imaging. These a…

eess.IV2026

SAMRI: Segment Any MRI

Zhao Wang, Wei Dai, Thuy Thanh Dao +4

Summary: SAMRI is an MRI-specialized adaptation of the Segment Anything Model achieving superior whole-body MRI segmentation, particularly for small and clinically critical structu…

cs.CV2025

CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects

Zhao Wang, Aoxue Li, Lingting Zhu +3

Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generatio…

cs.CV2025

AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance

Zhao Wang, Hao Wen, Lingting Zhu +3

Character video generation is a significant real-world application focused on producing high-quality videos featuring specific characters. Recent advancements have introduced vario…

cs.CV2025

Grounded Knowledge-Enhanced Medical Vision-Language Pre-training for Chest X-Ray

Qiao Deng, Zhongzhen Huang, Yunqi Wang +6

Medical foundation models have the potential to revolutionize healthcare by providing robust and generalized representations of medical data. Medical vision-language pre-training h…

cs.CV2025

Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting

Lingting Zhu, Guying Lin, Jinnan Chen +4

While Implicit Neural Representations (INRs) have demonstrated significant success in image representation, they are often hindered by large training memory and slow decoding speed…