1 paper · 1 filter
Hanling Yi, Feng Lin, Mao Luo +3
Recent advances in Multimodal Large Language Models (MLLMs) have enabled open-ended object recognition, yet they struggle with fine-grained tasks. In contrast, CLIP-style models ex…