activity
20192022
most citedTemporal Pyramid Transformer with Multimodal Interaction for Video Question Answering

5 citations · 10 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20231 cited

Efficient End-to-End Video Question Answering with Pyramidal Multimodal Transformer

Min Peng, Chongyang Wang, Yu Shi +1

This paper presents a new method for end-to-end Video Question Answering (VideoQA), aside from the current popularity of using large-scale pre-training with huge feature extractors…

cs.CV20225 cited

Multilevel Hierarchical Network with Multiscale Sampling for Video Question Answering

Min Peng, Chongyang Wang, Yuan Gao +2

Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language processing. While most existing approaches ignore th…

cs.CV20215 cited

Temporal Pyramid Transformer with Multimodal Interaction for Video Question Answering

Min Peng, Chongyang Wang, Yuan Gao +2

Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language understanding. While existing approaches seldom leve…

cs.CV2020

Recognizing Micro-Expression in Video Clip with Adaptive Key-Frame Mining

Min Peng, Chongyang Wang, Yuan Gao +4

As a spontaneous expression of emotion on face, micro-expression reveals the underlying emotion that cannot be controlled by human. In micro-expression, facial movement is transien…

cs.CV2019

A Novel Apex-Time Network for Cross-Dataset Micro-Expression Recognition

Min Peng, Chongyang Wang, Tao Bi +3

The automatic recognition of micro-expression has been boosted ever since the successful introduction of deep learning approaches. As researchers working on such topics are moving…