1 paper · 1 filter
Mingning Guo, Mengwei Wu, Shaoxian Li +2
Existing image perception methods based on VLMs generally follow a paradigm wherein models extract and analyze image content based on user-provided textual task prompts. However, s…