12 papers
Repurposing RGB-based Foundation Model for Depth Estimation on Thermal Images Using Hierarchical Supervision
Jie Hong, Tingtian Li, Xuesong Li +1
Depth estimation from thermal images is highly valuable for robotic applications in adverse conditions, such as nighttime and rainy weather. Recent studies have sought to transfer…
Keep the Future, Drop the Rollout: RIFT for World Action Models
Chushan Zhang, Jinguang Tong, Xuesong Li +2
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evo…
Maintain Plasticity in Long-timescale Continual Test-time Adaptation
Yanshuo Wang, Xuesong Li, Jinguang Tong +5
Continual test-time domain adaptation (CTTA) aims to adjust pre-trained source models to perform well over time across non-stationary target environments. While previous methods ha…
Structural Energy Guidance for View-Consistent Text-to-3D Generation
Qing Zhang, Jinguang Tong, Jing Zhang +2
Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work identifies viewpoint bias in 2D…
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
Chushan Zhang, Ruihan Lu, Jinguang Tong +3
Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet robot actions cause contact,…
EgoSelf: From Memory to Personalized Egocentric Assistant
Yanshuo Wang, Yuan Xu, Xuesong Li +4
Egocentric assistants often rely on first-person view data to capture user behavior and context for personalized services. Since different users exhibit distinct habits, preference…