3 papers
cs.CL2026
Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability
Yu-Ting Lee, Fu-Chieh Chang, Yu-En Shu +2
Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely through prompting. While effective across…
eess.AS2026
Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition
Bo-Hao Su, Hui-Ying Shih, Jinchuan Tian +4
Speech Emotion Recognition (SER) is typically trained and evaluated on majority-voted labels, which simplifies benchmarking but masks subjectivity and provides little transparency…
cs.AI2025
RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner
Fu-Chieh Chang, Yu-Ting Lee, Hui-Ying Shih +2
The reasoning abilities of large language models (LLMs) have improved with chain-of-thought (CoT) prompting, allowing models to solve complex tasks stepwise. However, training CoT…