9 papers
TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models
Mark Van der Merwe, Mohamad Louai Shehab, Jayjun Lee +4
Vision-Language-Action (VLA) models demonstrate impressive reasoning over visual, semantic, and spatial task variations by leveraging large-scale vision and language pre-training.…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
Yinpei Dai, Hongze Fu, Jayjun Lee +6
Memory is critical for long-horizon and history-dependent robotic manipulation. Such tasks often involve counting repeated actions or manipulating objects that become temporarily o…
HydroShear: Hydroelastic Shear Simulation for Tactile Sim-to-Real Reinforcement Learning
An Dang, Jayjun Lee, Mustafa Mukadam +4
In this paper, we address the problem of tactile sim-to-real policy transfer for contact-rich tasks. Existing methods primarily focus on vision-based sensors and emphasize image re…
Using Temperature Sampling to Effectively Train Robot Learning Policies on Imbalanced Datasets
Basavasagar Patil, Sydney Belt, Jayjun Lee +2
Increasingly large datasets of robot actions and sensory observations are being collected to train ever-larger neural networks. These datasets are collected based on tasks and whil…
Visual-auditory Extrinsic Contact Estimation
Xili Yi, Jayjun Lee, Nima Fazeli
Robust manipulation often hinges on a robot's ability to perceive extrinsic contacts-contacts between a grasped object and its surrounding environment. However, these contacts are…