activity
20242026
collaborators

14 papers

cs.IR2026

Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agents

Shuai Wang, Haodong Chen, Yu Yin +3

Existing deep-research agents use a Search--Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metad…

cs.IR2026

DiffRetriever: Parallel Representative Tokens for Retrieval with Diffusion Language Models

Shuai Wang, Yu Yin, Shengyao Zhuang +2

This paper shows how diffusion language models (DLMs) can be used as effective and efficient retrievers. Existing DLM-based retrievers (e.g., DiffEmbed) follow BERT-style encoding,…

cs.IR2026

Where Relevance Emerges: A Layer-Wise Study of Internal Attention for Zero-Shot Re-Ranking

Haodong Chen, Shengyao Zhuang, Zheng Yao +2

Zero-shot document re-ranking with Large Language Models (LLMs) has evolved from Pointwise methods to Listwise and Setwise approaches that optimize computational efficiency. Despit…

cs.IR2025

An Investigation of Prompt Variations for Zero-shot LLM-based Rankers

Shuoqi Sun, Shengyao Zhuang, Shuai Wang +1

We provide a systematic understanding of the impact of specific components and wordings used in prompts on the effectiveness of rankers based on zero-shot Large Language Models (LL…

cs.IR2025

MAGMaR Shared Task System Description: Video Retrieval with OmniEmbed

Jiaqi Samantha Zhan, Crystina Zhang, Shengyao Zhuang +2

Effective video retrieval remains challenging due to the complexity of integrating visual, auditory, and textual modalities. In this paper, we explore unified retrieval methods usi…

cs.IR2025

Starbucks-v2: Improved Training for 2D Matryoshka Embeddings

Shengyao Zhuang, Shuai Wang, Fabio Zheng +2

2D Matryoshka training enables a single embedding model to generate sub-network representations across different layers and embedding dimensions, offering adaptability to diverse c…