activity
20242026
collaborators

9 papers

cs.LG2026

Language-Critique Imitation Learning from Suboptimal Demonstrations

Chih-Han Yang, Dai-Jie Wu, Yun-Ping Huang +3

Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estimates, discriminator scores, or importance…

cs.LG2026

Plan2Cleanse: Test-Time Backdoor Defense via Monte-Carlo Planning in Deep Reinforcement Learning

Sze-Ann Chen, Zhi-Yi Chin, Kui-Yuan Chen +2

Ensuring the security of reinforcement learning (RL) models is critical, particularly when they are trained by third parties and deployed in real-world systems. Attackers can impla…

cs.CL2026

Test-Time Alignment for Large Language Models via Textual Model Predictive Control

Kuang-Da Wang, Teng-Ruei Chen, Yu Heng Hung +7

Aligning Large Language Models (LLMs) with human preferences through finetuning is resource-intensive, motivating lightweight alternatives at test time. We address test-time alignm…

cs.CL2025

Extending Automatic Machine Translation Evaluation to Book-Length Documents

Kuang-Da Wang, Shuoyang Ding, Chao-Han Huck Yang +4

Despite Large Language Models (LLMs) demonstrating superior translation performance and long-context capabilities, evaluation methodologies remain constrained to sentence-level ass…

cs.RO2025

Action-Constrained Imitation Learning

Chia-Han Yeh, Tse-Sheng Nan, Risto Vuorio +4

Policy learning under action constraints plays a central role in ensuring safe behaviors in various robot control and resource allocation applications. In this paper, we study a ne…

cs.AI2025

Imitation Learning of Correlated Policies in Stackelberg Games

Kuang-Da Wang, Ping-Chun Hsieh, Wen-Chih Peng

Stackelberg games, widely applied in domains like economics and security, involve asymmetric interactions where a leader's strategy drives follower responses. Accurately modeling t…