collaborators

6 papers

cs.CV2026

X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining

Miracle Kang, Lights Shi, Lucy Liang +10

Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing action tokenizers discretize actions prim…

cs.RO2026

Wall-OSS-0.5 Technical Report

Ryan Yu, Pushi Zhang, Starrick Liu +24

Large-scale Vision-Language-Action (VLA) pretraining is increasingly adopted as the foundation for robot policies, yet the evidence for pretrained VLAs is almost invariably reporte…

cs.LG2025

What Do Latent Action Models Actually Learn?

Chuheng Zhang, Tim Pearce, Pushi Zhang +5

Latent action models (LAMs) aim to learn action-relevant changes from unlabeled videos by compressing changes between frames as latents. However, differences between video frames c…

cs.RO2025

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models

Xiaoyu Chen, Hangxing Wei, Pushi Zhang +9

Vision-Language-Action (VLA) models have emerged as a popular paradigm for learning robot manipulation policies that can follow language instructions and generalize to novel scenar…

cs.CV2025

PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models

Jiansong Wan, Chengming Zhou, Jinkua Liu +14

Recent studies have explored pretrained (foundation) models for vision-based robotic navigation, aiming to achieve generalizable navigation and positive transfer across diverse env…

cs.CL2025

UC-MOA: Utility-Conditioned Multi-Objective Alignment for Distributional Pareto-Optimality

Zelei Cheng, Xin-Qiang Cai, Yuting Tang +4

Reinforcement Learning from Human Feedback (RLHF) has become a cornerstone for aligning large language models (LLMs) with human values. However, existing approaches struggle to cap…