7 papers
KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models
Xiao Zhang, Qianru Meng, Yongjian Chen +2
Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based…
Benchmark Leakage Trap: Can We Trust LLM-based Recommendation?
Mingqiao Zhang, Qiyao Peng, Yinghui Wang +2
The expanding integration of Large Language Models (LLMs) into recommender systems poses critical challenges to evaluation reliability. This paper identifies and investigates a pre…
RankSteer: Activation Steering for Pointwise LLM Ranking
Yumeng Wang, Catherine Chen, Suzan Verberne
Large language models (LLMs) have recently shown strong performance as zero-shot rankers, yet their effectiveness is highly sensitive to prompt formulation, particularly role-play…
How role-play shapes relevance judgment in zero-shot LLM rankers
Yumeng Wang, Jirui Qi, Catherine Chen +2
Large Language Models (LLMs) have emerged as promising zero-shot rankers, but their performance is highly sensitive to prompt formulation. In particular, role-play prompts, where t…
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
Yumeng Wang, Tianyu Fan, Lingrui Xu +1
Large Language Models (LLMs) have evolved from simple chatbots into sophisticated agents capable of automating complex real-world tasks, where browsing and reasoning over live web…
QUIDS: Query Intent Description for Exploratory Search via Dual Space Modeling
Yumeng Wang, Xiuying Chen, Suzan Verberne
In exploratory search, users often submit vague queries to investigate unfamiliar topics, but receive limited feedback about how the search engine understood their input. This lead…