1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.AI2026★ 1 cited
A unified multimodal understanding and generation model for cross-disciplinary scientific research
Xiaomeng Yang, Zhiyu Tan, Xiaohui Zhong +5
Scientific discovery increasingly relies on integrating heterogeneous, high-dimensional data across disciplines nowadays. While AI models have achieved notable success across vario…
cs.CV2025
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
Luozheng Qin, Zhiyu Tan, Mengping Yang +2
Video Detailed Captioning (VDC) is a crucial task for vision-language bridging, enabling fine-grained descriptions of complex video content. In this paper, we first comprehensively…
cs.CV2024
E2ED^2:Direct Mapping from Noise to Data for Enhanced Diffusion Models
Zhiyu Tan, WenXu Qian, Hesen Chen +3
Diffusion models have established themselves as the de facto primary paradigm in visual generative modeling, revolutionizing the field through remarkable success across various div…