4 papers
STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Stories
Qiuyu Tian, Zequn Liu, Yiding Li +10
Movie screenplays are a demanding testbed for long-form narrative understanding, as characters' goals, beliefs, knowledge, and relationships evolve continuously across scenes. Howe…
DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs
Bo-Cheng Chiu, Jen-Jee Chen, Yu-Chee Tseng +2
Large Language Models (LLMs) have recently been extended to the video domain, enabling sophisticated video-language understanding. However, existing Video LLMs often exhibit limita…
An Empirical Study on How Video-LLMs Answer Video Questions
Chenhui Gou, Ziyu Ma, Zicheng Duan +6
Taking advantage of large-scale data and pretrained language models, Video Large Language Models (Video-LLMs) have shown strong capabilities in answering video questions. However,…
MPT: Motion Prompt Tuning for Micro-Expression Recognition
Jiateng Liu, Hengcan Shi, Feng Chen +4
Micro-expression recognition (MER) is crucial in the affective computing field due to its wide application in medical diagnosis, lie detection, and criminal investigation. Despite…