3 papers
cs.CV2026
Dynamic Resolution Routing for Efficient Egocentric Grounding
Huixin Sun, Wangbo Zhao, Fanyue Wei +3
Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excess…
cs.CV2025
Semantics-aware Test-time Adaptation for 3D Human Pose Estimation
Qiuxia Lin, Rongyu Chen, Kerui Gu +1
This work highlights a semantics misalignment in 3D human pose estimation. For the task of test-time adaptation, the misalignment manifests as overly smoothed and unguided predicti…
cs.CV2025
Online Test-time Adaptation for 3D Human Pose Estimation: A Practical Perspective with Estimated 2D Poses
Qiuxia Lin, Kerui Gu, Linlin Yang +1
Online test-time adaptation for 3D human pose estimation is used for video streams that differ from training data. Ground truth 2D poses are used for adaptation, but only estimated…