1 paper
Yuzhong Zhao, Feng Liu, Yue Liu +4
One fundamental task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptabi…