39 citations · 81 across the 7 of their papers we have counts for
6 papers · 1 filter
Waver: Wave Your Way to Lifelike Video Generation
Yifu Zhang, Hao Yang, Yuqi Zhang +7
We present Waver, a high-performance foundation model for unified image and video generation. Waver can directly generate videos with durations ranging from 5 to 10 seconds at a na…
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
Chuang Lin, Bingbing Zhuang, Shanlin Sun +3
The recent advent of large-scale 3D data, e.g. Objaverse, has led to impressive progress in training pose-conditioned diffusion models for novel view synthesis. However, due to the…
Generative Region-Language Pretraining for Open-Ended Object Detection
Chuang Lin, Yi Jiang, Lizhen Qu +2
In recent research, significant attention has been devoted to the open-vocabulary object detection task, aiming to generalize beyond the limited number of classes labeled during tr…
Learning Object-Language Alignments for Open-Vocabulary Object Detection
Chuang Lin, Peize Sun, Yi Jiang +5
Existing object detection methods are bounded in a fixed-set vocabulary by costly labeled data. When dealing with novel categories, the model has to be retrained with more bounding…
Emotional Semantics-Preserved and Feature-Aligned CycleGAN for Visual Emotion Adaptation
Sicheng Zhao, Xuanbai Chen, Xiangyu Yue +7
Thanks to large-scale labeled training data, deep neural networks (DNNs) have obtained remarkable success in many vision and multimedia tasks. However, because of the presence of d…
Multi-source Domain Adaptation for Visual Sentiment Classification
Chuang Lin, Sicheng Zhao, Lei Meng +1
Existing domain adaptation methods on visual sentiment classification typically are investigated under the single-source scenario, where the knowledge learned from a source domain…