Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Dynamic Resolution Routing for Efficient Egocentric Grounding
Huixin Sun, Wangbo Zhao, Fanyue Wei +3
Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excess…
cs.CV2025
Visual Intention Grounding for Egocentric Assistants
Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse +3
Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as A…
cs.CV2025
Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose Estimation
Zhuoran Zhao, Linlin Yang, Pengzhan Sun +2
Recent synthetic 3D human datasets for the face, body, and hands have pushed the limits on photorealism. Face recognition and body pose estimation have achieved state-of-the-art pe…