4 papers
Backdooring Masked Diffusion Language Models
Daniel Yiming Cao, Chengzhong Wang, Sheng-Yen Chou +3
Masked diffusion language models (MDLMs) are emerging as a compelling new paradigm for text generation, but their training-time security remains largely unexplored. Existing backdo…
Token-weighted Direct Preference Optimization with Attention
Chengyu Huang, Zhuohang Li, Sheng-Yen Chou +1
Direct Preference Optimization (DPO) aligns Large Language Models with human preferences without the need for a separate reward model. However, DPO treats all tokens in responses e…
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text
Chengyu Huang, Sheng-Yen Chou, Zhengxin Zhang +1
Self-play has recently emerged as a promising paradigm for post-training Large Language Models (LLMs). In self-play, the target LLM creates the task input (e.g., a question), which…
Rethinking Backdoor Attacks on Dataset Distillation: A Kernel Method Perspective
Ming-Yu Chung, Sheng-Yen Chou, Chia-Mu Yu +3
Dataset distillation offers a potential means to enhance data efficiency in deep learning. Recent studies have shown its ability to counteract backdoor risks present in original tr…