4 papers
KV-Tracker: Real-Time Pose Tracking with Transformers
Marwan Taher, Ignacio Alzugaray, Kirill Mazur +2
Multi-view 3D geometry networks offer a powerful prior but are prohibitively slow for real-time applications. We propose a novel way to adapt them for online use, enabling real-tim…
STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
Feng Xu, Guangyao Zhai, Xin Kong +4
Recent advances in Vision-Language-Action (VLA) models, powered by large language models and reinforcement learning-based fine-tuning, have shown remarkable progress in robotic man…
CausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesis
Xin Kong, Daniel Watson, Yannick Strümpler +2
Multi-view diffusion models have shown promise in 3D novel view synthesis, but most existing methods adopt a non-autoregressive formulation. This limits their applicability in worl…
OminiAdapt: Learning Cross-Task Invariance for Robust and Environment-Aware Robotic Manipulation
Yongxu Wang, Weiyun Yi, Xinhao Kong +1
With the rapid development of embodied intelligence, leveraging large-scale human data for high-level imitation learning on humanoid robots has become a focal point of interest in…