most citedKALIE: Fine-Tuning Vision-Language Models for Open-World Manipulation without Robot Data

2 citations · 2 across the 4 of their papers we have counts for

collaborators

6 papers

cs.RO2025

Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins

Chuanruo Ning, Kuan Fang, Wei-Chiu Ma

Recent advancements in open-world robot manipulation have been largely driven by vision-language models (VLMs). While these models exhibit strong generalization ability in high-lev…

cs.RO2025

Towards Embodiment Scaling Laws in Robot Locomotion

Bo Ai, Liu Dai, Nico Bohlinger +7

Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodim…

cs.RO2025

Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following

Vivek Myers, Bill Chunyuan Zheng, Anca Dragan +2

Effective task representations should facilitate compositionality, such that after learning a variety of basic tasks, an agent can perform compound tasks consisting of multiple ste…

cs.RO2024

Blox-Net: Generative Design-for-Robot-Assembly Using VLM Supervision, Physics Simulation, and a Robot with Reset

Andrew Goldberg, Kavish Kondap, Tianshuang Qiu +7

Generative AI systems have shown impressive capabilities in creating text, code, and images. Inspired by the rich history of research in industrial ''Design for Assembly'', we intr…

cs.RO20242 cited

KALIE: Fine-Tuning Vision-Language Models for Open-World Manipulation without Robot Data

Grace Tang, Swetha Rajkumar, Yifei Zhou +3

Building generalist robotic systems involves effectively endowing robots with the capabilities to handle novel objects in an open-world setting. Inspired by the advances of large p…

cs.RO2024

Policy Adaptation via Language Optimization: Decomposing Tasks for Few-Shot Imitation

Vivek Myers, Bill Chunyuan Zheng, Oier Mees +2

Learned language-conditioned robot policies often struggle to effectively adapt to new real-world tasks even when pre-trained across a diverse set of instructions. We propose a nov…