6 papers
Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports
Haobin Li, Ping Deng, Weizhong Qian +4
Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase. Ho…
Reliable Thinking with Images
Haobin Li, Yutong Yang, Yijie Lin +3
As a multimodal extension of Chain-of-Thought (CoT), Thinking with Images (TWI) has recently emerged as a promising avenue to enhance the reasoning capability of Multi-modal Large…
ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge
Yijie Lin, Guofeng Ding, Haochen Zhou +3
Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To…
Toward Robust and Harmonious Adaptation for Cross-modal Retrieval
Haobin Li, Mouxing Yang, Xi Peng
Recently, the general-to-customized paradigm has emerged as the dominant approach for Cross-Modal Retrieval (CMR), which reconciles the distribution shift problem between the sourc…
Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment
Haobin Li, Yijie Lin, Peng Hu +2
Multi-modal entity alignment (MMEA) aims to identify equivalent entities across heterogeneous multi-modal knowledge graphs (MMKGs), where each entity is described by attributes fro…
Test-time Adaptation for Cross-modal Retrieval with Query Shift
Haobin Li, Peng Hu, Qianjun Zhang +3
The success of most existing cross-modal retrieval methods heavily relies on the assumption that the given queries follow the same distribution of the source domain. However, such…