2 citations · 4 across the 13 of their papers we have counts for
1 paper · 1 filter
Rang Li, Lei Li, Shuhuai Ren +10
Visual grounding, localizing objects from natural language descriptions, represents a critical bridge between language and vision understanding. While multimodal large language mod…