From the 1 of 4 linked papers with an AI index.
4 papers
Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
The paper introduces a system-level acceleration for vision-language-action models by incrementally updating visual tokens for dynamic regions and compressing diffusion-based polic…
RoBoSR: Structured Scene Representations for Embodied Robotic Reasoning
Kewei Hu, Wanchan Yu, Fangwen Chen +6
Despite rapid progress, embodied reasoning under real-world variability remains challenging. Existing approaches rely on demonstration-driven sequential biases, limiting flexibilit…
OMNI-PoseX: A Fast Vision Model for 6D Object Pose Estimation in Embodied Tasks
Michael Zhang, Wei Ying, Fangwen Chen +2
Accurate 6D object pose estimation is a fundamental capability for embodied agents, yet remains highly challenging in open-world environments. Many existing methods often rely on c…
GSR: Learning Structured Reasoning for Embodied Manipulation
Kewei Hu, Michael Zhang, Wei Ying +7
Despite rapid progress, embodied agents still struggle with long-horizon manipulation that requires maintaining spatial consistency, causal dependencies, and goal constraints. A ke…