collaborators

8 papers

cs.CV2026

PROBE: Manipulation-Grounded Visual Question Answering with VLM Agents

Vineet Bhat, Siyi Chen, Alex Zook +4

Vision-language Models (VLMs) excel at 2D grounding, spatial reasoning and agentic tool-based planning in static scenes. However, consider asking a home robot "Is my medication sti…

cs.CV2026

BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

Vineet Bhat, Sungsu Kim, Valts Blukis +6

Vision Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses in understanding object interacti…

cs.AI2026

RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers

Vineet Bhat, Shiqing Wei, Ali Umut Kaypak +3

Reconstructing numerical simulations from control systems research papers is often hindered by underspecified parameters and ambiguous implementation details. We define the task of…

cs.RO2026

3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks

Vineet Bhat, Yu-Hsiang Lan, Prashanth Krishnamurthy +2

Robotic manipulation in 3D requires effective computation of N degree-of-freedom joint-space trajectories that enable precise and robust control. To achieve this, robots must integ…

cs.RO2025

Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback

Vineet Bhat, Ali Umut Kaypak, Prashanth Krishnamurthy +2

Planning algorithms decompose complex problems into intermediate steps that can be sequentially executed by robots to complete tasks. Recent works have employed Large Language Mode…

cs.RO2025

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping

Vineet Bhat, Naman Patel, Prashanth Krishnamurthy +2

Robotic manipulation of unseen objects via natural language commands remains challenging. Language driven robotic grasping (LDRG) predicts stable grasp poses from natural language…