6 papers
Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen +15
AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which…
Game-Time: Evaluating Temporal Dynamics in Spoken Language Models
Kai-Wei Chang, En-Pei Hu, Chun-Yi Kuan +7
Conversational Spoken Language Models (SLMs) are emerging as a promising paradigm for real-time speech interaction. However, their capacity of temporal dynamics, including the abil…
On Calibration of Large Language Models: From Response To Capability
Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin +3
Large language models (LLMs) are widely deployed as general-purpose problem solvers, making accurate confidence estimation critical for reliable use. Prior work on LLM calibration…
Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
Ren-Wei Liang, Chin-Ting Hsu, Chan-Hung Yu +6
Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessive refusals, while permissive mode…
Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations
Li-Chun Lu, Miri Liu, Pin-Chun Lu +3
We examine, analyze, and compare four representative creativity measures--perplexity, LLM-as-a-Judge, the Creativity Index (CI; measuring n-gram overlap with web corpora), and synt…
BILLY: Steering Large Language Models via Merging Persona Vectors for Creative Generation
Tsung-Min Pai, Jui-I Wang, Li-Chun Lu +3
Multi-LLM systems enhance the creativity of large language models by simulating human collective intelligence but suffer from significant drawbacks, such as high computational cost…