1 citations · 1 across the 6 of their papers we have counts for
6 papers
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
Zhefan Rao, Liya Ji, Yazhou Xing +6
Text-to-video (T2V) generation has gained significant attention recently. However, the costs of training a T2V model from scratch remain persistently high, and there is considerabl…
Large Motion Video Autoencoding with Cross-modal Video VAE
Yazhou Xing, Yang Fei, Yingqing He +4
Learning a robust video Variational Autoencoder (VAE) is essential for reducing video redundancy and facilitating efficient video generation. Directly applying image VAEs to indivi…
High-fidelity 3D GAN Inversion by Pseudo-multi-view Optimization
Jiaxin Xie, Hao Ouyang, Jingtan Piao +2
We present a high-fidelity 3D generative adversarial network (GAN) inversion framework that can synthesize photo-realistic novel views while preserving specific details of the inpu…
Dual-Camera Super-Resolution with Aligned Attention Modules
Tengfei Wang, Jiaxin Xie, Wenxiu Sun +2
We present a novel approach to reference-based super-resolution (RefSR) with the focus on dual-camera super-resolution (DCSR), which utilizes reference images for high-quality and…
Depth Sensing Beyond LiDAR Range
Kai Zhang, Jiaxin Xie, Noah Snavely +1
Depth sensing is a critical component of autonomous driving technologies, but today's LiDAR- or stereo camera-based solutions have limited range. We seek to increase the maximum ra…
Video Depth Estimation by Fusing Flow-to-Depth Proposals
Jiaxin Xie, Chenyang Lei, Zhuwen Li +2
Depth from a monocular video can enable billions of devices and robots with a single camera to see the world in 3D. In this paper, we present an approach with a differentiable flow…