works on

From the 2 of 11 linked papers with an AI index.

collaborators

11 papers

cs.CV2026

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA

Ping-Kun Chiang, Kun-Ru Wu, Po-han Li +3

ViewMind3D is a training‑free, modular framework that answers 3D questions by selecting relevant views, grounding objects with language cues, encoding spatial context via a bird's‑…

cs.CV2026

VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR

Shenghui Chen, Po-han Li, Sandeep Chinchali +1

The paper introduces VIBE, an annotation-free method that evaluates video-to-text summaries by measuring how well they match visual content and how useful they are for downstream d…

cs.CV2026

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Shenghui Chen, Po-han Li, Ximeng Sun +5

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation…

cs.AI2026

What We are Missing in Multimodal LLM Evaluation?

Po-han Li, Shenghui Chen, Sandeep Chinchali +1

Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced ra…

cs.CV2026

SSR: A Generic Framework for Text-Aided Map Compression for Localization

Mohammad Omama, Po-han Li, Harsh Goel +6

Mapping is crucial in robotics for localization and downstream decision-making. As robots are deployed in ever-broader settings, the maps they rely on continue to increase in size.…

cs.CV2026

ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning

Po-han Li, Shenghui Chen, Ufuk Topcu +1

Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors gen…