4 papers · 1 filter
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
Vineet Bhat, Yu-Hsiang Lan, Prashanth Krishnamurthy +2
Robotic manipulation in 3D requires effective computation of N degree-of-freedom joint-space trajectories that enable precise and robust control. To achieve this, robots must integ…
Grounding LLMs For Robot Task Planning Using Closed-loop State Feedback
Vineet Bhat, Ali Umut Kaypak, Prashanth Krishnamurthy +2
Planning algorithms decompose complex problems into intermediate steps that can be sequentially executed by robots to complete tasks. Recent works have employed Large Language Mode…
MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping
Vineet Bhat, Naman Patel, Prashanth Krishnamurthy +2
Robotic manipulation of unseen objects via natural language commands remains challenging. Language driven robotic grasping (LDRG) predicts stable grasp poses from natural language…
HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models
Vineet Bhat, Prashanth Krishnamurthy, Ramesh Karri +1
Robots interacting with humans through natural language can unlock numerous applications such as Referring Grasp Synthesis (RGS). Given a text query, RGS determines a stable grasp…