10 papers
Can LLM Rerankers Predict Their Own Ranking Performance?
Shiyu Ni, Keping Bi, Jiafeng Guo +3
Retrieval effectiveness varies substantially across queries, making it important to estimate ranking quality before relevance judgments are available. Query performance prediction…
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
Yuhan Wang, Shiyu Ni, Zhikai Ding +3
Confidence calibration is essential for making large language models (LLMs) reliable, yet existing training-free methods have been primarily studied under single-answer question an…
How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality
Minzhu Tu, Shiyu Ni, Keping Bi
Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One pos…
Annotation-Efficient Universal Honesty Alignment
Shiyu Ni, Keping Bi, Jiafeng Guo +4
Honesty alignment-the ability of large language models (LLMs) to recognize their knowledge boundaries and express calibrated confidence-is essential for trustworthy deployment. Exi…
Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
Minghao Tang, Shiyu Ni, Jingtong Wu +2
Context-grounded generation underpins many LLM applications, including long-document question answering (QA), conversational personalization, and retrieval-augmented generation (RA…
Deep Research: A Systematic Survey
Zhengliang Shi, Yiqun Chen, Haitao Li +23
Large language models (LLMs) have rapidly evolved from text generators into powerful problem solvers. Yet, many open tasks demand critical thinking, multi-source, and verifiable ou…