1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.AI2026★ 1 cited
A unified multimodal understanding and generation model for cross-disciplinary scientific research
Xiaomeng Yang, Zhiyu Tan, Xiaohui Zhong +5
Scientific discovery increasingly relies on integrating heterogeneous, high-dimensional data across disciplines nowadays. While AI models have achieved notable success across vario…
cs.CV2024★ 1 cited
VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
Zhiyu Tan, Xiaomeng Yang, Luozheng Qin +1
The quality of video-text pairs fundamentally determines the upper bound of text-to-video models. Currently, the datasets used for training these models suffer from significant sho…
cs.CV2024
Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image Generation
Junyan Wang, Zhenhong Sun, Zhiyu Tan +5
Vanilla text-to-image diffusion models struggle with generating accurate human images, commonly resulting in imperfect anatomies such as unnatural postures or disproportionate limb…