9 papers
PRISM: A Geometric Risk Bound that Decomposes Drift into Scale, Shape, and Head
Chieh-Yen Lin, Shao-Hua Sun
Comparing post-training LLM variants, such as quantized, LoRA-adapted, and distilled models, requires a diagnostic that identifies how a variant has drifted, not only whether it ha…
On Calibration of Large Language Models: From Response To Capability
Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin +3
Large language models (LLMs) are widely deployed as general-purpose problem solvers, making accurate confidence estimation critical for reliable use. Prior work on LLM calibration…
Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning
Chao-Chung Wu, Zhi Rui Tam, Chieh-Yen Lin +3
Maintaining consistent model performance across domains is a fundamental challenge in machine learning. While recent work has explored using LLM-generated data for fine-tuning, its…
Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models
Cheng-Kuang Wu, Zhi Rui Tam, Chieh-Yen Lin +2
Language models (LMs) are increasingly used to build agents that can act autonomously to achieve goals. During this automatic process, agents need to take a series of actions, some…
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
Zhi Rui Tam, Cheng-Kuang Wu, Yu Ying Chiu +3
Large reasoning models (LRMs) have demonstrated impressive performance across a range of reasoning tasks, yet little is known about their internal reasoning processes in multilingu…
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering
Zhi Rui Tam, Cheng-Kuang Wu, Chieh-Yen Lin +1
Multiple-choice exam questions with "None of the above" (NA) options have been extensively studied in educational testing, in which existing research suggests that they better asse…