collaborators

6 papers

cs.IR2026

Sparton: Fast and Memory-Efficient Triton Kernel for Learned Sparse Retrieval

Thong Nguyen, Cosimo Rulli, Franco Maria Nardini +2

State-of-the-art Learned Sparse Retrieval (LSR) models, such as Splade, typically employ a Language Modeling (LM) head to project latent hidden states into a lexically-anchored log…

cs.IR2026

Milco: Learned Sparse Retrieval Across Languages via a Multilingual Connector

Thong Nguyen, Yibin Lei, Jia-Huei Ju +2

Learned Sparse Retrieval (LSR) combines the efficiency of bi-encoders with the transparency of lexical matching, but existing approaches struggle to scale beyond English. We introd…

cs.CL2026

Enhancing Lexicon-Based Text Embeddings with Large Language Models

Yibin Lei, Tao Shen, Yu Cao +1

Recent large language models (LLMs) have demonstrated exceptional performance on general-purpose text embedding tasks. While dense embeddings have dominated related research, we in…

cs.IR2026

ThinkQE: Query Expansion via an Evolving Thinking Process

Yibin Lei, Tao Shen, Andrew Yates

Effective query expansion for web search benefits from promoting both exploration and result diversity to capture multiple interpretations and facets of a query. While recent LLM-b…

cs.IR2025

Making Large Language Models Efficient Dense Retrievers

Yibin Lei, Shwai He, Ang Li +1

Recent work has shown that directly fine-tuning large language models (LLMs) for dense retrieval yields strong performance, but their substantial parameter counts make them computa…

cs.IR2025

SERVAL: Surprisingly Effective Zero-Shot Visual Document Retrieval Powered by Large Vision and Language Models

Thong Nguyen, Yibin Lei, Jia-Huei Ju +1

Visual Document Retrieval (VDR) typically operates as text-to-image retrieval using specialized bi-encoders trained to directly embed document images. We revisit a zero-shot genera…