3 papers
cs.CV2025
EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models
Yufei Cai, Hu Han, Yuxiang Wei +2
The progress on generative models has led to significant advances on text-to-video (T2V) generation, yet the motion controllability of generated videos remains limited. Existing mo…
cs.CV2025
Dynamically evolving segment anything model with continuous learning for medical image segmentation
Zhaori Liu, Mengyang Li, Hu Han +3
Medical image segmentation is essential for clinical diagnosis, surgical planning, and treatment monitoring. Traditional approaches typically strive to tackle all medical image seg…
cs.CV2024
Face-MLLM: A Large Face Perception Model
Haomiao Sun, Mingjie He, Tianheng Lian +2
Although multimodal large language models (MLLMs) have achieved promising results on a wide range of vision-language tasks, their ability to perceive and understand human faces is…