From the 1 of 7 linked papers with an AI index.
7 papers
Unified Visuomotor Targets: Supervising VLAs Beyond Physical Actions
Zhenyang Feng, Unnat Jain
VLA models are trained to predict robot actions from visual and language observations. This is a natural choice, but it creates a mismatch: VLMs encode rich, high-level representat…
DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
Mengqi Zhang, Sahil Khose, Simar Kareer +3
Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models be…
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
Dwip Dalal, Shivansh Patel, Chahit Jain +7
The paper introduces Anchor-Align, a method that adds representation anchoring and language-action alignment to behavior‑cloning finetuning of vision‑language models for robot mani…
CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance
Leo Lin, Shivansh Patel, Jay Moon +2
We introduce CRAFT hand, a tendon-driven anthropomorphic hand with hybrid hard-soft compliance for contact-rich manipulation. The design is based on a simple idea: contact is not u…
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
Dwip Dalal, Gautam Vashishtha, Utkarsh Mishra +6
Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce At…
ViPRA: Video Prediction for Robot Actions
Sandeep Routray, Hengkai Pan, Unnat Jain +2
Can we turn a video prediction model into a robot policy? Videos, including those of humans or teleoperated robots, capture rich physical interactions. However, most of them lack l…