activity
20242026
collaborators
Showing cs.ROShow all

18 papers · 1 filter

cs.RO2026

VANDERER: Map-Free Exploration using Future-Aware and Visual-Curiosity-Guided Diffusion Policy

Venkata Naren Devarakonda, Raktim Gautam Goswami, Prashanth Krishnamurthy +1

Mobile agents require efficient exploration strategies to map unseen environments and autonomously plan tasks. Traditional methods rely on generating occupancy maps and optimizing…

cs.RO2026

Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks

Raktim Gautam Goswami, Prashanth Krishnamurthy, Yann LeCun +1

Visual world models have shown great potential in learning complex system dynamics. Recent advancements leverage these models as transition functions within Model Predictive Contro…

cs.RO2026

Open-Architecture End-to-End System for Real-World Autonomous Robot Navigation

Venkata Naren Devarakonda, Ali Umut Kaypak, Raktim Gautam Goswami +4

Enabling robots to autonomously navigate unknown, complex, and dynamic real-world environments presents several challenges, including imperfect perception, partial observability, l…

cs.RO2026

3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks

Vineet Bhat, Yu-Hsiang Lan, Prashanth Krishnamurthy +2

Robotic manipulation in 3D requires effective computation of N degree-of-freedom joint-space trajectories that enable precise and robust control. To achieve this, robots must integ…

cs.RO2026

World Models for Learning Dexterous Hand-Object Interactions from Human Videos

Raktim Gautam Goswami, Amir Bar, David Fan +6

Modeling dexterous hand-object interactions is challenging as it requires understanding how subtle finger motions influence the environment through contact with objects. While rece…

cs.RO2025

OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory Generation

Raktim Gautam Goswami, Prashanth Krishnamurthy, Yann LeCun +1

Visual imitation learning enables robotic agents to acquire skills by observing expert demonstration videos. In the one-shot setting, the agent generates a policy after observing a…