37 citations · 47 across the 10 of their papers we have counts for
9 papers · 1 filter
Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation
Xiaochuan Li, Guoguang Du, Runze Zhang +11
Scaling laws have validated the success and promise of large-data-trained models in creative generation across text, image, and video domains. However, this paradigm faces data sca…
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
Runze Zhang, Guoguang Du, Xiaochuan Li +10
Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visu…
First Place Solution to the ECCV 2024 ROAD++ Challenge @ ROAD++ Spatiotemporal Agent Detection 2024
Tengfei Zhang, Heng Zhang, Ruyang Li +3
This report presents our team's solutions for the Track 1 of the 2024 ECCV ROAD++ Challenge. The task of Track 1 is spatiotemporal agent detection, which aims to construct an "agen…
Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline
Qi Jia, Baoyu Fan, Cong Xu +7
Existing video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watch…
Image Content Generation with Causal Reasoning
Xiaochuan Li, Baoyu Fan, Runze Zhang +5
The emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed…
Real-Aug: Realistic Scene Synthesis for LiDAR Augmentation in 3D Object Detection
Jinglin Zhan, Tiejun Liu, Rengang Li +3
Data and model are the undoubtable two supporting pillars for LiDAR object detection. However, data-centric works have fallen far behind compared with the ever-growing list of fanc…