3 papers
cs.CL2026
ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions
Chuanyang Jin, Binze Li, Haopeng Xie +6
Conversational AI has now reached billions of users, yet existing datasets capture only what people say, not what they think. We introduce ThoughtTrace, the first large-scale datas…
cs.CL2026
I-CALM: Incentivizing Confidence-Aware Abstention for LLM Selective Answering
Haotian Zong, Binze Li, Yufei Long +3
Large language models (LLMs) often produce confident but incorrect answers, in part because standard evaluation incentives reward guessing over expressing uncertainty. We study epi…
cs.AI2025
A Study of Rule Omission in Raven's Progressive Matrices
Binze Li
Analogical reasoning lies at the core of human cognition and remains a fundamental challenge for artificial intelligence. Raven's Progressive Matrices (RPM) serve as a widely used…