4 papers · 1 filter
Unveiling the Latent Directions of Reflection in Large Language Models
Fu-Chieh Chang, Yu-Ting Lee, Pei-Yuan Wu
Reflection, the ability of large language models (LLMs) to evaluate and revise their own reasoning, has been widely used to improve performance on complex reasoning tasks. Yet, mos…
A Theoretical Framework for OOD Robustness in Transformers using Gevrey Classes
Yu Wang, Fu-Chieh Chang, Pei-Yuan Wu
We study the robustness of Transformer language models under semantic out-of-distribution (OOD) shifts, where training and test data lie in disjoint latent spaces. Using Wasserstei…
Unraveling Arithmetic in Large Language Models: The Role of Algebraic Structures
Fu-Chieh Chang, You-Chen Lin, Pei-Yuan Wu
Large language models (LLMs) have demonstrated remarkable mathematical capabilities, largely driven by chain-of-thought (CoT) prompting, which decomposes complex reasoning into ste…
Leveraging Unlabeled Data Sharing through Kernel Function Approximation in Offline Reinforcement Learning
Yen-Ru Lai, Fu-Chieh Chang, Pei-Yuan Wu
Offline reinforcement learning (RL) learns policies from a fixed dataset, but often requires large amounts of data. The challenge arises when labeled datasets are expensive, especi…