7 papers
Backdooring Masked Diffusion Language Models
Daniel Yiming Cao, Chengzhong Wang, Sheng-Yen Chou +3
Masked diffusion language models (MDLMs) are emerging as a compelling new paradigm for text generation, but their training-time security remains largely unexplored. Existing backdo…
Token-weighted Direct Preference Optimization with Attention
Chengyu Huang, Zhuohang Li, Sheng-Yen Chou +1
Direct Preference Optimization (DPO) aligns Large Language Models with human preferences without the need for a separate reward model. However, DPO treats all tokens in responses e…
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
Chengyu Huang, Sheng-Yen Chou, Zhengxin Zhang +1
Self-play has recently emerged as a promising paradigm for post-training Large Language Models (LLMs). In self-play, the target LLM creates the task input (e.g., a question), which…
Better LLM Reasoning via Dual-Play
Zhengxin Zhang, Chengyu Huang, Aochong Oliver Li +1
Large Language Models (LLMs) have achieved remarkable progress through Reinforcement Learning with Verifiable Rewards (RLVR), yet still rely heavily on external supervision (e.g.,…
HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
Chengyu Huang, Zhengxin Zhang, Claire Cardie
While scaling the length of responses at test-time has been shown to markedly improve the reasoning abilities and performance of large language models (LLMs), it often results in v…
Judging with Confidence: Calibrating Autoraters to Preference Distributions
Zhuohang Li, Xiaowei Li, Chengyu Huang +11
The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limite…