activity
20182022
most citedZooming Slow-Mo: Fast and Accurate One-Stage Space-Time Video Super-Resolution

20 citations · 50 across the 9 of their papers we have counts for

collaborators

16 papers

cs.CV20225 cited

Learning to Answer Questions in Dynamic Audio-Visual Scenarios

Guangyao Li, Yake Wei, Yapeng Tian +3

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in vid…

eess.IV20223 cited

Transformer-empowered Multi-scale Contextual Matching and Aggregation for Multi-contrast MRI Super-resolution

Guangyuan Li, Jun Lv, Yapeng Tian +4

Magnetic resonance imaging (MRI) can present multi-contrast images of the same anatomical structures, enabling multi-contrast super-resolution (SR) techniques. Compared with SR rec…

cs.CV2022

Learning Spatio-Temporal Downsampling for Effective Video Upscaling

Xiaoyu Xiang, Yapeng Tian, Vijay Rengarajan +3

Downsampling is one of the most basic image processing operations. Improper spatio-temporal downsampling applied on videos can cause aliasing issues such as moiré patterns in space…

cs.CV20217 cited

Zooming SlowMo: An Efficient One-Stage Framework for Space-Time Video Super-Resolution

Xiaoyu Xiang, Yapeng Tian, Yulun Zhang +3

In this paper, we address the space-time video super-resolution, which aims at generating a high-resolution (HR) slow-motion video from a low-resolution (LR) and low frame rate (LF…

cs.CV20213 cited

Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation

Yapeng Tian, Di Hu, Chenliang Xu

There are rich synchronized audio and visual events in our daily life. Inside the events, audio scenes are associated with the corresponding visual objects; meanwhile, sounding obj…

cs.CV20211 cited

Can audio-visual integration strengthen robustness under multimodal attacks?

Yapeng Tian, Chenliang Xu

In this paper, we propose to make a systematic study on machines multisensory perception under attacks. We use the audio-visual event recognition task against multimodal adversaria…