works on

From the 4 of 112 linked papers with an AI index.

activity
20242026
most citedRandomized Greedy Methods for Weak Submodular Sensor Selection with Robustness Considerations

3 citations · 3 across the 32 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

If, Then, Otherwise: Diagnosing Conditional Branching in Vision-Language Navigation

Seoyoung Lee, Neel P. Bhatt, Pranay Samineni +8

Vision-language navigation agents are often evaluated on their ability to follow route-like instructions toward a fixed goal. Yet, real navigation instructions often depend on obse…

cs.CV2026

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA

Ping-Kun Chiang, Kun-Ru Wu, Po-han Li +3

ViewMind3D is a training‑free, modular framework that answers 3D questions by selecting relevant views, grounding objects with language cues, encoding spatial context via a bird's‑…

cs.CV2026

VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR

Shenghui Chen, Po-han Li, Sandeep Chinchali +1

The paper introduces VIBE, an annotation-free method that evaluates video-to-text summaries by measuring how well they match visual content and how useful they are for downstream d…

cs.CV2026

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Shenghui Chen, Po-han Li, Ximeng Sun +5

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation…

cs.CV2026

ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning

Po-han Li, Shenghui Chen, Ufuk Topcu +1

Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors gen…

cs.CV2025

Real-Time Privacy Preservation for Robot Visual Perception

Minkyu Choi, Yunhao Yang, Neel P. Bhatt +6

Many robots (e.g., iRobot's Roomba) operate based on visual observations from live video streams, and such observations may inadvertently include privacy-sensitive objects, such as…