4 papers
Querit-Reranker: Training Compact Multilingual Rerankers via Efficient Label-Free Distribution Adaptation
Yunfei Zhong, Jun Yang, Wei Huang +7
Deployable multilingual rerankers must generalize across languages, domains, and target ranking tasks while remaining efficient enough for second-stage reranking. However, adapting…
Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning
Jiahan Chen, Da Li, Hengran Zhang +6
Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classi…
Data, Not Model: Explaining Bias toward LLM Texts in Neural Retrievers
Wei Huang, Keping Bi, Yinqiong Cai +3
Recent studies show that neural retrievers often display source bias, favoring passages generated by LLMs over human-written ones, even when both are semantically similar. This bia…
How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
Wei Huang, Keping Bi, Yinqiong Cai +3
As more content generated by large language models (LLMs) floods into the Internet, information retrieval (IR) systems now face the challenge of distinguishing and handling a blend…