1 citations · 2 across the 11 of their papers we have counts for
Showing 2024Show all
3 papers · 1 filter
cs.CV2024
LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment
Yibin Wang, Zhiyu Tan, Junyan Wang +3
Recent advances in text-to-video (T2V) generative models have shown impressive capabilities. However, these models are still inadequate in aligning synthesized videos with human pr…
cs.CV2024★ 1 cited
VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
Zhiyu Tan, Xiaomeng Yang, Luozheng Qin +1
The quality of video-text pairs fundamentally determines the upper bound of text-to-video models. Currently, the datasets used for training these models suffer from significant sho…
cs.CV2024
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
Zhiyu Tan, Xiaomeng Yang, Luozheng Qin +3
The recent advancements in text-to-image generative models have been remarkable. Yet, the field suffers from a lack of evaluation metrics that accurately reflect the performance of…