37 citations · 40 across the 6 of their papers we have counts for
8 papers
Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation
Xiaochuan Li, Guoguang Du, Runze Zhang +11
Scaling laws have validated the success and promise of large-data-trained models in creative generation across text, image, and video domains. However, this paradigm faces data sca…
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
Runze Zhang, Guoguang Du, Xiaochuan Li +10
Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visu…
Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline
Qi Jia, Baoyu Fan, Cong Xu +7
Existing video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watch…
Image Content Generation with Causal Reasoning
Xiaochuan Li, Baoyu Fan, Runze Zhang +5
The emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed…
UniMAE: Multi-modal Masked Autoencoders with Unified 3D Representation for 3D Perception in Autonomous Driving
Jian Zou, Tianyu Huang, Guanglei Yang +4
Masked Autoencoders (MAE) play a pivotal role in learning potent representations, delivering outstanding results across various 3D perception tasks essential for autonomous driving…
Meta-Auxiliary Network for 3D GAN Inversion
Bangrui Jiang, Zhenhua Guo, Yujiu Yang
Real-world image manipulation has achieved fantastic progress in recent years. GAN inversion, which aims to map the real image to the latent code faithfully, is the first step in t…