5 papers · 1 filter
Robust Multi-view Clustering against Imperfect Information
Zhichao Huang, Haochen Zhou, Hao Wang +2
Real-world multi-view data always suffer from imperfect information problem, where the view-specific observations are absent (i.e., Incomplete Views, IV) and cross-view corresponde…
Reliable Thinking with Images
Haobin Li, Yutong Yang, Yijie Lin +3
As a multimodal extension of Chain-of-Thought (CoT), Thinking with Images (TWI) has recently emerged as a promising avenue to enhance the reasoning capability of Multi-modal Large…
ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge
Yijie Lin, Guofeng Ding, Haochen Zhou +3
Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To…
LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification
Yiding Lu, Mouxing Yang, Dezhong Peng +3
Traditional text-based person ReID assumes that person descriptions from witnesses are complete and provided at once. However, in real-world scenarios, such descriptions are often…
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
Yuxin Tian, Mouxing Yang, Yunfan Li +4
Recent studies applied Parameter Efficient Fine-Tuning techniques (PEFTs) to efficiently narrow the performance gap between pre-training and downstream. There are two important fac…