From the 1 of 8 linked papers with an AI index.
8 papers
VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
Shenghui Chen, Po-han Li, Sandeep Chinchali +1
The paper introduces VIBE, an annotation-free method that evaluates video-to-text summaries by measuring how well they match visual content and how useful they are for downstream d…
VEGAS: Human-Aligned Video Caption Evaluation via Gaze
Shenghui Chen, Po-han Li, Ximeng Sun +5
Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation…
What We are Missing in Multimodal LLM Evaluation?
Po-han Li, Shenghui Chen, Sandeep Chinchali +1
Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced ra…
ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
Po-han Li, Shenghui Chen, Ufuk Topcu +1
Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors gen…
IG-MCTS: Human-in-the-Loop Cooperative Navigation under Incomplete Information
Shenghui Chen, Ruihan Zhao, Sandeep Chinchali +1
Human-robot cooperative navigation is challenging under incomplete information. We introduce CoNav-Maze, a simulated environment where a robot navigates with local perception while…
Learning to Coordinate without Communication under Incomplete Information
Shenghui Chen, Shufang Zhu, Giuseppe De Giacomo +1
Achieving seamless coordination in cooperative games is a crucial challenge in artificial intelligence, particularly when players operate under incomplete information. While commun…