From the 1 of 37 linked papers with an AI index.
37 papers
Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination
Harsh Goel, Aditya Sai Ellendula, Vaishnav Tadiparthi +3
Multi-agent Large Language Model (LLM) systems often struggle to collaborate with new teammates whose strategies shift mid-task. Because agents execute multi-step or temporally ext…
If, Then, Otherwise: Diagnosing Conditional Branching in Vision-Language Navigation
Seoyoung Lee, Neel P. Bhatt, Pranay Samineni +8
Vision-language navigation agents are often evaluated on their ability to follow route-like instructions toward a fixed goal. Yet, real navigation instructions often depend on obse…
ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA
Ping-Kun Chiang, Kun-Ru Wu, Po-han Li +3
ViewMind3D is a training‑free, modular framework that answers 3D questions by selecting relevant views, grounding objects with language cues, encoding spatial context via a bird's‑…
VEGAS: Human-Aligned Video Caption Evaluation via Gaze
Shenghui Chen, Po-han Li, Ximeng Sun +5
Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation…
Incentivizing Vision Language Models to Search for Long Video Question Answering
Harsh Goel, S P Sharan, Sahil Shah +4
We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek…
What We are Missing in Multimodal LLM Evaluation?
Po-han Li, Shenghui Chen, Sandeep Chinchali +1
Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses. While their capabilities have advanced ra…