3 papers
cs.CV2026
EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
Zexuan Yan, Yuzhou Wu, Yue Ma +9
Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We pre…
cs.RO2026
Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary so…
cs.CV2025
IPCV: Information-Preserving Compression for MLLM Visual Encoders
Yuan Chen, Zichen Wen, Yuzhou Wu +6
Multimodal Large Language Models (MLLMs) deliver strong vision-language performance but at high computational cost, driven by numerous visual tokens processed by the Vision Transfo…