5 papers
An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
Zhi Li, Yanan Wang, Hao Niu +2
Multimodal large language models have recently achieved remarkable progress in video question answering (VideoQA) by jointly processing visual, textual, and audio information. Howe…
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
Yanan Wang, Julio Vizcarra, Zhi Li +2
Despite recent progress in video large language models (VideoLLMs), a key open challenge remains: how to equip models with chain-of-thought (CoT) reasoning abilities grounded in fi…
Zero-shot Persuasive Chatbots with LLM-Generated Strategies and Information Retrieval
Kazuaki Furumai, Roberto Legaspi, Julio Vizcarra +6
Persuasion plays a pivotal role in a wide range of applications from health intervention to the promotion of social good. Persuasive chatbots employed responsibly for social good c…
Top-down Activity Representation Learning for Video Question Answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng +2
Capturing complex hierarchical human activities, from atomic actions (e.g., picking up one present, moving to the sofa, unwrapping the present) to contextual events (e.g., celebrat…
Multi-object event graph representation learning for Video Question Answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng +2
Video question answering (VideoQA) is a task to predict the correct answer to questions posed about a given video. The system must comprehend spatial and temporal relationships amo…