18 papers
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
Yubo Huang, Hailong Guo, Fangtai Wu +9
Audio-driven avatar interaction demands real-time, streaming, and infinite-length generation -- capabilities fundamentally at odds with the sequential denoising and long-horizon dr…
A Reasoning-Enabled Vision-Language Foundation Model for Chest X-ray Interpretation
Yabin Zhang, Chong Wang, Yunhe Gao +19
Chest X-rays (CXRs) are among the most frequently performed imaging examinations worldwide, yet rising imaging volumes increase radiologist workload and the risk of diagnostic erro…
The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment
Hongyuan Liu, Qinli Yang, Wen Li +6
Vision-Language Models (VLMs) such as CLIP learn a shared embedding space for images and text, yet their representations remain geometrically separated, a phenomenon known as the m…
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
Yulin Luo, Hao Chen, Zhuangzhe Wu +10
Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for robotic manipulation, in which reliable action prediction critically depends on accurately int…
Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs
Yiren Zheng, Shibo Li, Jiaming Liu +2
Current multimodal approaches predominantly treat visual generation as an external process, relying on pixel rendering or code execution, thereby overlooking the native visual repr…
Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
Yunhe Gao, Yabin Zhang, Chong Wang +5
Foundation models have transformed vision and language by learning general-purpose representations from large-scale unlabeled data, yet 3D medical imaging lacks analogous approache…