activity
20242026
collaborators

8 papers

cs.RO2026

WSA: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control

Jiahao Jiang, Jianing Zhang, Zhenhan Yin +8

Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for generalist robotic systems to date. By leveraging imitation learning on…

cs.CV2026

Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization

Ke Liu, Xuanhan Wang, Qilong Zhang +2

Deep image watermarking, which refers to enabling imperceptible watermark embedding and reliable extraction in cover images, has been shown to be effective for copyright protection…

cs.RO2025

MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training

Zhenhan Yin, Xuanhan Wang, Jiahao Jiang +8

While leveraging abundant human videos and simulated robot data poses a scalable solution to the scarcity of real-world robot data, the generalization capability of existing vision…

cs.CV2025

Pseudo-Label Refinement for Robust Wheat Head Segmentation via Two-Stage Hybrid Training

Jiahao Jiang, Zhangrui Yang, Xuanhan Wang +1

This extended abstract details our solution for the Global Wheat Full Semantic Segmentation Competition. We developed a systematic self-training framework. This framework combines…

cs.CV2025

Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models

Xuanhan Wang, Huimin Deng, Ke Liu +3

Human-centric vision models (HVMs) have achieved remarkable generalization due to large-scale pretraining on massive person images. However, their dependence on large neural archit…

cs.CV2025

Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models

Xuanhan Wang, Huimin Deng, Lianli Gao +1

Human-centric visual perception (HVP) has recently achieved remarkable progress due to advancements in large-scale self-supervised pretraining (SSP). However, existing HVP models f…