activity
20182022
most citedWord-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods Comparison

53 citations · 201 across the 23 of their papers we have counts for

collaborators
Showing 2021 · cs.CVShow all

13 papers · 2 filters

cs.CV2021

One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation Learning

Suzhen Wang, Lincheng Li, Yu Ding +1

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes a…

cs.CV2021★ 8 cited

RGB-D Saliency Detection via Cascaded Mutual Information Minimization

Jing Zhang, Deng-Ping Fan, Yuchao Dai +4

Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multi-stage cascad…

cs.CV2021

PR-RRN: Pairwise-Regularized Residual-Recursive Networks for Non-rigid Structure-from-Motion

Haitian Zeng, Yuchao Dai, Xin Yu +2

We propose PR-RRN, a novel neural-network based method for Non-rigid Structure-from-Motion (NRSfM). PR-RRN consists of Residual-Recursive Networks (RRN) and two extra regularizatio…

cs.CV2021★ 2 cited

VidFace: A Full-Transformer Solver for Video FaceHallucination with Unaligned Tiny Snapshots

Yuan Gan, Yawei Luo, Xin Yu +2

In this paper, we investigate the task of hallucinating an authentic high-resolution (HR) human face from multiple low-resolution (LR) video snapshots. We propose a pure transforme…

cs.CV2021★ 36 cited

VTNet: Visual Transformer Network for Object Goal Navigation

Heming Du, Xin Yu, Liang Zheng

Object goal navigation aims to steer an agent towards a target object based on observations of the agent. It is of pivotal importance to design effective visual representations of…

cs.CV2021★ 8 cited

Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation

Lincheng Li, Suzhen Wang, Zhimeng Zhang +4

In this paper, we propose a novel text-based talking-head video generation framework that synthesizes high-fidelity facial expressions and head motions in accordance with contextua…