6 papers · 1 filter
Keep the Future, Drop the Rollout: RIFT for World Action Models
Chushan Zhang, Jinguang Tong, Xuesong Li +2
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evo…
Repurposing RGB-based Foundation Model for Depth Estimation on Thermal Images Using Hierarchical Supervision
Jie Hong, Tingtian Li, Xuesong Li +1
Depth estimation from thermal images is highly valuable for robotic applications in adverse conditions, such as nighttime and rainy weather. Recent studies have sought to transfer…
Structural Energy Guidance for View-Consistent Text-to-3D Generation
Qing Zhang, Jinguang Tong, Jing Zhang +2
Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work identifies viewpoint bias in 2D…
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
Chushan Zhang, Ruihan Lu, Jinguang Tong +3
Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet robot actions cause contact,…
EgoSelf: From Memory to Personalized Egocentric Assistant
Yanshuo Wang, Yuan Xu, Xuesong Li +4
Egocentric assistants often rely on first-person view data to capture user behavior and context for personalized services. Since different users exhibit distinct habits, preference…
Adaptive and Balanced Re-initialization for Long-timescale Continual Test-time Domain Adaptation
Yanshuo Wang, Jinguang Tong, Jun Lan +5
Continual test-time domain adaptation (CTTA) aims to adjust models so that they can perform well over time across non-stationary environments. While previous methods have made cons…