2 papers
cs.CV2023
Is a Video worth Images? A Highly Efficient Approach to Transformer-based Video Question Answering
Chenyang Lyu, Tianbo Ji, Yvette Graham +1
Conventional Transformer-based Video Question Answering (VideoQA) approaches generally encode frames independently through one or more image encoders followed by interaction betwee…
cs.CV2023
Semantic-aware Dynamic Retrospective-Prospective Reasoning for Event-level Video Question Answering
Chenyang Lyu, Tianbo Ji, Yvette Graham +1
Event-Level Video Question Answering (EVQA) requires complex reasoning across video events to obtain the visual information needed to provide optimal answers. However, despite sign…