activity
20242026
collaborators

7 papers

cs.LG2026

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone

DatologyAI, :, Siddharth Joshi +32

Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language models (VLMs) is far less establi…

cs.RO2025

Towards Embodiment Scaling Laws in Robot Locomotion

Bo Ai, Liu Dai, Nico Bohlinger +7

Cross-embodiment generalization underpins the vision of building generalist embodied agents for any robot, yet its enabling factors remain poorly understood. We investigate embodim…

cs.RO2025

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…

cs.CV2025

FlexCap: Describe Anything in Images in Controllable Detail

Debidatta Dwibedi, Vidhi Jain, Jonathan Tompson +2

We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths. FlexCap is trained to produce length-conditioned captions for input bo…

cs.RO2024

RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning

Charles Xu, Qiyang Li, Jianlan Luo +1

Recent advances in robotic foundation models have enabled the development of generalist policies that can adapt to diverse tasks. While these models show impressive flexibility, th…

cs.RO2024

ANAVI: Audio Noise Awareness using Visuals of Indoor environments for NAVIgation

Vidhi Jain, Rishi Veerapaneni, Yonatan Bisk

We propose Audio Noise Awareness using Visuals of Indoors for NAVIgation for quieter robot path planning. While humans are naturally aware of the noise they make and its impact on…