activity
20232026
most citedA Comparative Study of Training Objectives for Clarification Facet Generation

2 citations · 3 across the 28 of their papers we have counts for

collaborators
Showing cs.CLShow all

13 papers · 1 filter

cs.CL2026

How Long Reasoning Chains Influence LLMs' Judgment of Answer Factuality

Minzhu Tu, Shiyu Ni, Keping Bi

Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation, yet such judges remain imperfect and susceptible to surface-level biases. One pos…

cs.CL2026

Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers

Yuhan Wang, Shiyu Ni, Zhikai Ding +3

Confidence calibration is essential for making large language models (LLMs) reliable, yet existing training-free methods have been primarily studied under single-answer question an…

cs.CL2025

Annotation-Efficient Universal Honesty Alignment

Shiyu Ni, Keping Bi, Jiafeng Guo +4

Honesty alignment-the ability of large language models (LLMs) to recognize their knowledge boundaries and express calibrated confidence-is essential for trustworthy deployment. Exi…

cs.CL2025

LLM-Specific Utility for Retrieval-Augmented Generation

Hengran Zhang, Keping Bi, Jiafeng Guo +4

Retrieval-augmented generation (RAG) is typically optimized for topical relevance, yet its success ultimately depends on whether retrieved passages are useful for a large language…

cs.CL2025

Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMs

Zhikai Ding, Shiyu Ni, Keping Bi

Large vision-language models (LVLMs) demonstrate strong visual question answering (VQA) capabilities but are shown to hallucinate. A reliable model should perceive its knowledge bo…

cs.CL20251 cited

Popular but Wrong: Understanding and Mitigating LLM Overconfidence through Knowledge Popularity

Shiyu Ni, Keping Bi, Jiafeng Guo +1

Large language models (LLMs) often produce incorrect answers with high confidence, yet the factors associated with such overconfidence remain insufficiently understood. We study th…