11 papers
Not All Layers Need Tuning: Diagnosing and Directing Adaptation in Vision-Language-Action Models
Shahram Najam Syed, Arthur Jakobsson, Prayuj Sachdev +1
Fine-tuning a Vision-Language-Action (VLA) model for a new deployment environment is expensive, yet most methods apply uniform-capacity adapters to every network region as if every…
Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models
Haoran Hao, Shahram Najam Syed, Jeff Schneider +1
While Vision-Language-Action (VLA) models pretrained on large-scale robot datasets provide a strong foundation for robot manipulation, their performance can degrade when adapted to…
FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement
Haoran Hao, Shahram Najam Syed, Jeffrey Ichnowski +1
Robot policies inevitably encounter failures when deployed in real environments. Naive retries often repeat the same mistakes, while many existing recovery methods rely on human in…
Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation
Shahram Najam Syed, Arthur Jakobsson, Haoran Hao +1
Vision-Language-Action (VLA) models generalize across static manipulation but fail when objects move during task execution. They map the current observation to an action and assume…
Wiggle and Go! System Identification for Zero-Shot Dynamic Rope Manipulation
Arthur Jakobsson, Abhinav Mahajan, Karthik Pullalarevu +6
Many robotic tasks are unforgiving; a single mistake in a dynamic throw can lead to unacceptable delays or unrecoverable failure. We introduce Wiggle and Go!, a two-stage framework…
Joint Learning of Depth, Pose, and Local Radiance Field for Large Scale Monocular 3D Reconstruction
Shahram Najam Syed, Yitian Hu, Yuchao Yao
Photorealistic 3-D reconstruction from monocular video collapses in large-scale scenes when depth, pose, and radiance are solved in isolation: scale-ambiguous depth yields ghost ge…