activity
20172021
most citedStructure-Preserving Image Super-resolution via Contextualized Multi-task Learning

62 citations · 139 across the 10 of their papers we have counts for

collaborators

13 papers

cs.CV20219 cited

Temporal Contrastive Graph Learning for Video Action Recognition and Retrieval

Yang Liu, Keze Wang, Haoyuan Lan +1

Attempt to fully discover the temporal diversity and chronological characteristics for self-supervised video representation learning, this work takes advantage of the temporal depe…

cs.CV20205 cited

Knowledge-Routed Visual Question Reasoning: Challenges for Deep Representation Embedding

Qingxing Cao, Bailin Li, Xiaodan Liang +2

Though beneficial for encouraging the Visual Question Answering (VQA) models to discover the underlying knowledge by exploiting the input-output correlation beyond image and text c…

cs.RO20203 cited

Continuous Transition: Improving Sample Efficiency for Continuous Control Problems via MixUp

Junfan Lin, Zhongzhan Huang, Keze Wang +3

Although deep reinforcement learning (RL) has been successfully applied to a variety of robotic control tasks, it's still challenging to apply it to real-world tasks, due to the po…

cs.CV2020

Linguistically Driven Graph Capsule Network for Visual Question Reasoning

Qingxing Cao, Xiaodan Liang, Keze Wang +1

Recently, studies of visual question answering have explored various architectures of end-to-end networks and achieved promising results on both natural and synthetic datasets, whi…

cs.CV201936 cited

Instance-Aware Representation Learning and Association for Online Multi-Person Tracking

Hefeng Wu, Yafei Hu, Keze Wang +3

Multi-Person Tracking (MPT) is often addressed within the detection-to-association paradigm. In such approaches, human detections are first extracted in every frame and person traj…

cs.CV2019

Face Hallucination by Attentive Sequence Optimization with Reinforcement Learning

Yukai Shi, Guanbin Li, Qingxing Cao +2

Face hallucination is a domain-specific super-resolution problem that aims to generate a high-resolution (HR) face image from a low-resolution~(LR) input. In contrast to the existi…