activity
20172022
most citedTricorNet: A Hybrid Temporal Convolutional and Recurrent Network for Video Action Segmentation

55 citations · 191 across the 25 of their papers we have counts for

collaborators

39 papers

cs.CV20225 cited

Learning to Answer Questions in Dynamic Audio-Visual Scenarios

Guangyao Li, Yake Wei, Yapeng Tian +3

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in vid…

cs.CV20222 cited

StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis

Zhiheng Li, Martin Renqiang Min, Kai Li +1

Although progress has been made for text-to-image synthesis, previous methods fall short of generalizing to unseen or underrepresented attribute compositions in the input text. Lac…

eess.IV20223 cited

Transformer-empowered Multi-scale Contextual Matching and Aggregation for Multi-contrast MRI Super-resolution

Guangyuan Li, Jun Lv, Yapeng Tian +4

Magnetic resonance imaging (MRI) can present multi-contrast images of the same anatomical structures, enabling multi-contrast super-resolution (SR) techniques. Compared with SR rec…

cs.CV20221 cited

Cross-modal Contrastive Distillation for Instructional Activity Anticipation

Zhengyuan Yang, Jingen Liu, Jing Huang +4

In this study, we aim to predict the plausible future action steps given an observation of the past and study the task of instructional activity anticipation. Unlike previous antic…

cs.CV2021

Procedure Planning in Instructional Videos via Contextual Modeling and Model-based Policy Learning

Jing Bi, Jiebo Luo, Chenliang Xu

Learning new skills by observing humans' behaviors is an essential capability of AI. In this work, we leverage instructional videos to study humans' decision-making processes, focu…

cs.CV2021

Learning to Generate Scene Graph from Natural Language Supervision

Yiwu Zhong, Jing Shi, Jianwei Yang +2

Learning from image-text data has demonstrated recent success for many recognition tasks, yet is currently limited to visual features or individual visual concepts such as objects.…