Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
Yik Lung Pang, Changjae Oh
Given a textual description, the task of referring expression comprehension (REC) involves the localisation of the referred object in an image. Multimodal large language models (ML…
cs.CV2024
Sparse multi-view hand-object reconstruction for unseen environments
Yik Lung Pang, Changjae Oh, Andrea Cavallaro
Recent works in hand-object reconstruction mainly focus on the single-view and dense multi-view settings. On the one hand, single-view methods can leverage learned shape priors to…