collaborators
Showing cs.ROShow all

5 papers · 1 filter

cs.RO2026

Unified Visuomotor Targets: Supervising VLAs Beyond Physical Actions

Zhenyang Feng, Unnat Jain

VLA models are trained to predict robot actions from visual and language observations. This is a natural choice, but it creates a mismatch: VLMs encode rich, high-level representat…

cs.RO2026

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

Mengqi Zhang, Sahil Khose, Simar Kareer +3

Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models be…

cs.RO2026

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

Dwip Dalal, Shivansh Patel, Chahit Jain +7

Finetuning a pretrained vision-language model (VLM) on robot demonstrations via behavior cloning (BC) has become the standard recipe for vision-language-action (VLA) policies. Howe…

cs.RO2026

CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance

Leo Lin, Shivansh Patel, Jay Moon +2

We introduce CRAFT hand, a tendon-driven anthropomorphic hand with hybrid hard-soft compliance for contact-rich manipulation. The design is based on a simple idea: contact is not u…

cs.RO2025

ViPRA: Video Prediction for Robot Actions

Sandeep Routray, Hengkai Pan, Unnat Jain +2

Can we turn a video prediction model into a robot policy? Videos, including those of humans or teleoperated robots, capture rich physical interactions. However, most of them lack l…