7 papers
UniTacVLA: Unified Tactile Understanding and Prediction in Vision Language Action Models
Xidong Zhang, Yichi Zhang, Jiaxin Shi +5
Vision-language-action (VLA) models have achieved strong performance in many robotic manipulation tasks, yet remain limited in contact-rich dexterous manipulation. To overcome this…
LACE: Latent Visual Representation for Cross-Embodiment Learning
Yoo Sung Jang, Kanchana Ranasinghe, Cristina Mata +3
Cross-embodiment learning from human demonstrations is hindered by the visual gap between human and robot embodiments. While self-supervised learning (SSL) backbones encode rich in…
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
Yichi Zhang, Weihao Yuan, Yizhuo Zhang +2
Vision-Language-Action (VLA) models improve action generation by conditioning policies on rich vision-language information. However, current auto-regressive policies are constraine…
Evaluating Long-Context Reasoning in LLM-Based WebAgents
Andy Chung, Yichi Zhang, Kaixiang Lin +3
As large language model (LLM)-based agents become increasingly integrated into daily digital interactions, their ability to reason across long interaction histories becomes crucial…
AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies
Yinpei Dai, Jayjun Lee, Yichi Zhang +6
In this paper, we propose AimBot, a lightweight visual augmentation technique that provides explicit spatial cues to improve visuomotor policy learning in robotic manipulation. Aim…
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
Yichi Zhang, Xin Luna Dong, Zhaojiang Lin +5
Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactiv…