3 citations · 3 across the 3 of their papers we have counts for
4 papers · 1 filter
VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
Xiangdong Zhang, Jiaqi Liao, Shaofeng Zhang +4
Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physica…
PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders
Xiangdong Zhang, Shaofeng Zhang, Junchi Yan
Masked autoencoder has been widely explored in point cloud self-supervised learning, whereby the point cloud is generally divided into visible and masked parts. These methods typic…
Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based Approach
Shaofeng Zhang, Jinfa Huang, Qiang Zhou +4
Image outpainting aims to generate the content of an input sub-image beyond its original boundaries. It is an important task in content generation yet remains an open problem for g…
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
Ning Liao, Shaofeng Zhang, Renqiu Xia +3
There is an emerging line of research on multimodal instruction tuning, and a line of benchmarks has been proposed for evaluating these models recently. Instead of evaluating the m…