From the 2 of 11 linked papers with an AI index.
11 papers
ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA
Ping-Kun Chiang, Kun-Ru Wu, Po-han Li +3
ViewMind3D is a training‑free, modular framework that answers 3D questions by selecting relevant views, grounding objects with language cues, encoding spatial context via a bird's‑…
VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
Shenghui Chen, Po-han Li, Sandeep Chinchali +1
The paper introduces VIBE, an annotation-free method that evaluates video-to-text summaries by measuring how well they match visual content and how useful they are for downstream d…
VEGAS: Human-Aligned Video Caption Evaluation via Gaze
Shenghui Chen, Po-han Li, Ximeng Sun +5
Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation…
What We are Missing in Multimodal LLM Evaluation?
Po-han Li, Shenghui Chen, Sandeep Chinchali +1
Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced ra…
SSR: A Generic Framework for Text-Aided Map Compression for Localization
Mohammad Omama, Po-han Li, Harsh Goel +6
Mapping is crucial in robotics for localization and downstream decision-making. As robots are deployed in ever-broader settings, the maps they rely on continue to increase in size.…
ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
Po-han Li, Shenghui Chen, Ufuk Topcu +1
Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors gen…