5 citations · 5 across the 1 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023★ 1 cited
Efficient End-to-End Video Question Answering with Pyramidal Multimodal Transformer
Min Peng, Chongyang Wang, Yu Shi +1
This paper presents a new method for end-to-end Video Question Answering (VideoQA), aside from the current popularity of using large-scale pre-training with huge feature extractors…
cs.CV2022★ 5 cited
Multilevel Hierarchical Network with Multiscale Sampling for Video Question Answering
Min Peng, Chongyang Wang, Yuan Gao +2
Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language processing. While most existing approaches ignore th…