2 papers
cs.CV2026
Dynamic Resolution Routing for Efficient Egocentric Grounding
Huixin Sun, Wangbo Zhao, Fanyue Wei +3
Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excess…
cs.CV2025
Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose Estimation
Zhuoran Zhao, Linlin Yang, Pengzhan Sun +2
Recent synthetic 3D human datasets for the face, body, and hands have pushed the limits on photorealism. Face recognition and body pose estimation have achieved state-of-the-art pe…