2 papers
cs.CV2024
Top-down Activity Representation Learning for Video Question Answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng +2
Capturing complex hierarchical human activities, from atomic actions (e.g., picking up one present, moving to the sofa, unwrapping the present) to contextual events (e.g., celebrat…
cs.CV2024
Multi-object event graph representation learning for Video Question Answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng +2
Video question answering (VideoQA) is a task to predict the correct answer to questions posed about a given video. The system must comprehend spatial and temporal relationships amo…