59 citations · 113 across the 14 of their papers we have counts for
23 papers · 1 filter
Learning Semantic Facial Descriptors for Accurate Face Animation
Lei Zhu, Yuanqi Chen, Xiaohang Liu +2
Face animation is a challenging task. Existing model-based methods (utilizing 3DMMs or landmarks) often result in a model-like reconstruction effect, which doesn't effectively pres…
EPContrast: Effective Point-level Contrastive Learning for Large-scale Point Cloud Understanding
Zhiyi Pan, Guoqing Liu, Wei Gao +1
The acquisition of inductive bias through point-level contrastive learning holds paramount significance in point cloud pre-training. However, the square growth in computational req…
Mug-STAN: Adapting Image-Language Pretrained Models for General Video Understanding
Ruyang Liu, Jingjia Huang, Wei Gao +2
Large-scale image-language pretrained models, e.g., CLIP, have demonstrated remarkable proficiency in acquiring general multi-modal knowledge through web-scale image-text data. Des…
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation
Peihao Chen, Dongyu Ji, Kunyang Lin +4
We address a practical yet challenging problem of training robot agents to navigate in an environment following a path described by some language instructions. The instructions oft…
Learning Active Camera for Multi-Object Navigation
Peihao Chen, Dongyu Ji, Kunyang Lin +5
Getting robots to navigate to multiple objects autonomously is essential yet difficult in robot applications. One of the key challenges is how to explore environments efficiently w…
Frequency-Aware Self-Supervised Monocular Depth Estimation
Xingyu Chen, Thomas H. Li, Ruonan Zhang +1
We present two versatile methods to generally enhance self-supervised monocular depth estimation (MDE) models. The high generalizability of our methods is achieved by solving the f…