8 papers
Quality over Quantity: Demonstration Curation via Influence Functions for Data-Centric Robot Learning
Haeone Lee, Taywon Min, Junsu Kim +4
Learning from demonstrations has emerged as a promising paradigm for end-to-end robot control, particularly when scaled to diverse and large datasets. However, the quality of demon…
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
Huang Huang, Fangchen Liu, Letian Fu +5
Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained visio…
ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface
Fangchen Liu, Chuanyu Li, Yihua Qin +3
Tactile information plays a crucial role for humans and robots to interact effectively with their environment, particularly for tasks requiring the understanding of contact propert…
Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos
Weirui Ye, Fangchen Liu, Zheng Ding +3
Simulation offers a promising approach for cheaply scaling training data for generalist policies. To scalably generate data from diverse and realistic tasks, existing algorithms ei…
In-Context Imitation Learning via Next-Token Prediction
Letian Fu, Huang Huang, Gaurav Datta +5
We explore how to enhance next-token prediction models to perform in-context imitation learning on a real robot, where the robot executes new tasks by interpreting contextual infor…
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting
Fangchen Liu, Kuan Fang, Pieter Abbeel +1
Open-world generalization requires robotic systems to have a profound understanding of the physical world and the user command to solve diverse and complex tasks. While the recent…