4 papers
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
Jiale Zhao, Ke Fang, Lu Cheng
Large language models (LLMs) often respond even when prompts omit critical details or include misleading information, leading to hallucinations or reinforced misconceptions. We stu…
REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models
Ke Fang, Tianyi Zhao, Lu Cheng +1
Large language models (LLMs) often express verbal confidence that is poorly aligned with actual correctness, limiting their reliability in safety-critical applications. Existing pr…
Evaluating internal and external dissonance of belief dynamics in social systems
Joshua T. S. Hewson, Ke Fang
Belief dynamics are fundamental to human behavior and social coordination. Individuals rely on accurate beliefs to make decisions, and shared beliefs form the basis of successful c…
TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles
Qingchen Yu, Shichao Song, Ke Fang +5
As the application of Large Language Models (LLMs) expands, the demand for reliable evaluations increases. Existing LLM evaluation benchmarks primarily rely on static datasets, mak…