works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.RO2026

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

Ian Chuang, Jinyu Zou, Andrew Lee +2

The paper introduces GIAVA, a robot vision system that incorporates human-like gaze and foveated processing into Vision Transformers, reducing computational load and improving robu…

cs.CV2026

VITA: Vision-to-Action Flow Matching Policy

Dechen Gao, Boqi Zhao, Andrew Lee +6

Conventional flow matching and diffusion-based policies sample via iterative denoising from standard noise distributions (e.g., Gaussian), and require conditioning modules to repea…

cs.CV2025

Gaze on the Prize: Shaping Visual Attention with Return-Guided Contrastive Learning

Andrew Lee, Ian Chuang, Dechen Gao +2

Visual Reinforcement Learning (RL) agents must learn to act based on high-dimensional image data where only a small fraction of the pixels is task-relevant. This forces agents to w…

cs.RO2025

Mechanistic interpretability for steering vision-language-action models

Bear Häon, Kaylene Stocking, Ian Chuang +1

Vision-Language-Action (VLA) models are a promising path to realizing generalist embodied agents that can quickly adapt to new tasks, modalities, and environments. However, methods…

cs.RO2025

Active Vision Might Be All You Need: Exploring Active Vision in Bimanual Robotic Manipulation

Ian Chuang, Andrew Lee, Dechen Gao +2

Imitation learning has demonstrated significant potential in performing high-precision manipulation tasks using visual feedback. However, it is common practice in imitation learnin…

cs.RO2024

InterACT: Inter-dependency Aware Action Chunking with Hierarchical Attention Transformers for Bimanual Manipulation

Andrew Lee, Ian Chuang, Ling-Yuan Chen +1

Bimanual manipulation presents unique challenges compared to unimanual tasks due to the complexity of coordinating two robotic arms. In this paper, we introduce InterACT: Inter-dep…