1 paper · 1 filter
Yijie Lin, Guofeng Ding, Haochen Zhou +3
Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To…