works on

From the 1 of 43 linked papers with an AI index.

activity
20242026
most citedR3DM: Enabling Role Discovery and Diversity Through Dynamics Models in Multi-agent Reinforcement Learning

1 citations · 1 across the 8 of their papers we have counts for

collaborators

43 papers

cs.MA2026

Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination

Harsh Goel, Aditya Sai Ellendula, Vaishnav Tadiparthi +3

Multi-agent Large Language Model (LLM) systems often struggle to collaborate with new teammates whose strategies shift mid-task. Because agents execute multi-step or temporally ext…

cs.CV2026

If, Then, Otherwise: Diagnosing Conditional Branching in Vision-Language Navigation

Seoyoung Lee, Neel P. Bhatt, Pranay Samineni +8

Vision-language navigation agents are often evaluated on their ability to follow route-like instructions toward a fixed goal. Yet, real navigation instructions often depend on obse…

cs.CV2026

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA

Ping-Kun Chiang, Kun-Ru Wu, Po-han Li +3

ViewMind3D is a training‑free, modular framework that answers 3D questions by selecting relevant views, grounding objects with language cues, encoding spatial context via a bird's‑…

cs.RO2026

RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perception

Oguzhan Baser, Mirac Sozen, Kaan Kale +2

With the increased adoption of robotic agents operating in human environments by scanning and sharing 3D representations (e.g., for fleet learning, cloud-based planning, or collabo…

cs.CV2026

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Shenghui Chen, Po-han Li, Ximeng Sun +5

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation…

cs.CV2026

Incentivizing Vision Language Models to Search for Long Video Question Answering

Harsh Goel, S P Sharan, Sahil Shah +4

We introduce VSeek, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process. VSeek…