1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.AI2026★ 1 cited
A unified multimodal understanding and generation model for cross-disciplinary scientific research
Xiaomeng Yang, Zhiyu Tan, Xiaohui Zhong +5
Scientific discovery increasingly relies on integrating heterogeneous, high-dimensional data across disciplines nowadays. While AI models have achieved notable success across vario…
cs.CV2025
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
Luozheng Qin, Zhiyu Tan, Mengping Yang +2
Video Detailed Captioning (VDC) is a crucial task for vision-language bridging, enabling fine-grained descriptions of complex video content. In this paper, we first comprehensively…
cs.CV2024
LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment
Yibin Wang, Zhiyu Tan, Junyan Wang +3
Recent advances in text-to-video (T2V) generative models have shown impressive capabilities. However, these models are still inadequate in aligning synthesized videos with human pr…