5 papers
Reliable Thinking with Images
Haobin Li, Yutong Yang, Yijie Lin +3
As a multimodal extension of Chain-of-Thought (CoT), Thinking with Images (TWI) has recently emerged as a promising avenue to enhance the reasoning capability of Multi-modal Large…
ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge
Yijie Lin, Guofeng Ding, Haochen Zhou +3
Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To…
Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment
Haobin Li, Yijie Lin, Peng Hu +2
Multi-modal entity alignment (MMEA) aims to identify equivalent entities across heterogeneous multi-modal knowledge graphs (MMKGs), where each entity is described by attributes fro…
Incomplete Multi-view Clustering via Diffusion Contrastive Generation
Yuanyang Zhang, Yijie Lin, Weiqing Yan +6
Incomplete multi-view clustering (IMVC) has garnered increasing attention in recent years due to the common issue of missing data in multi-view datasets. The primary approach to ad…
LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification
Yiding Lu, Mouxing Yang, Dezhong Peng +3
Traditional text-based person ReID assumes that person descriptions from witnesses are complete and provided at once. However, in real-world scenarios, such descriptions are often…