3 papers
cs.CV2026
EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
Zexuan Yan, Yuzhou Wu, Yue Ma +9
The paper introduces EgoGenesis, a simulator that generates controllable egocentric manipulation videos using geometry-aware conditioning mechanisms to augment real robot data and…
cs.RO2026
Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
The paper introduces a system-level acceleration for vision-language-action models by incrementally updating visual tokens for dynamic regions and compressing diffusion-based polic…
cs.CV2025
IPCV: Information-Preserving Compression for MLLM Visual Encoders
Yuan Chen, Zichen Wen, Yuzhou Wu +6
Multimodal Large Language Models (MLLMs) deliver strong vision-language performance but at high computational cost, driven by numerous visual tokens processed by the Vision Transfo…