activity
20242026
collaborators

9 papers

cs.CL2026

On Calibration of Large Language Models: From Response To Capability

Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin +3

Large language models (LLMs) are widely deployed as general-purpose problem solvers, making accurate confidence estimation critical for reliable use. Prior work on LLM calibration…

cs.CR2026

Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs

Yen-Shan Chen, Zhi Rui Tam, Cheng-Kuang Wu +1

Current evaluations of LLM safety predominantly rely on severity-based taxonomies to assess the harmfulness of malicious queries. We argue that this formulation requires re-examina…

cs.CL2025

Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models

Cheng-Kuang Wu, Zhi Rui Tam, Chieh-Yen Lin +2

Language models (LMs) are increasingly used to build agents that can act autonomously to achieve goals. During this automatic process, agents need to take a series of actions, some…

cs.CL2025

Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?

Zhi Rui Tam, Cheng-Kuang Wu, Yu Ying Chiu +3

Large reasoning models (LRMs) have demonstrated impressive performance across a range of reasoning tasks, yet little is known about their internal reasoning processes in multilingu…

cs.CY2025

None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering

Zhi Rui Tam, Cheng-Kuang Wu, Chieh-Yen Lin +1

Multiple-choice exam questions with "None of the above" (NA) options have been extensively studied in educational testing, in which existing research suggests that they better asse…

cs.CL2024

StreamBench: Towards Benchmarking Continuous Improvement of Language Agents

Cheng-Kuang Wu, Zhi Rui Tam, Chieh-Yen Lin +2

Recent works have shown that large language model (LLM) agents are able to improve themselves from experience, which is an important ability for continuous enhancement post-deploym…