2 papers
cs.CV2025
HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
Lei Hu, Yongjing Ye, Shihong Xia
The expansion of instruction-tuning data has enabled foundation language models to exhibit improved instruction adherence and superior performance across diverse downstream tasks.…
cs.CV2025
Learning Transformation-Isomorphic Latent Space for Accurate Hand Pose Estimation
Kaiwen Ren, Lei Hu, Zhiheng Zhang +2
Vision-based regression tasks, such as hand pose estimation, have achieved higher accuracy and faster convergence through representation learning. However, existing representation…