collaborators

5 papers

cs.LG2026

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Xiaohua Wang, Muzhao Tian, Yuqi Zeng +20

Reinforcement Learning from Human Feedback (RLHF) and related alignment paradigms have become central to steering large language models (LLMs) and multimodal large language models…

cs.DC2025

SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading

Yuanzhe Shen, Yide Liu, Zisu Huang +3

Large language models (LLMs) demonstrate remarkable performance across diverse tasks, yet their effectiveness frequently depends on costly commercial APIs or cloud services. Model…

cs.AI2025

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement

Yuanzhe Shen, Zisu Huang, Zhengkang Guo +5

The rapid advancement of large language models (LLMs) has driven their adoption across diverse domains, yet their ability to generate harmful content poses significant safety chall…

cs.CL2025

Improving Continual Pre-training Through Seamless Data Packing

Ruicheng Yin, Xuan Gao, Changze Lv +3

Continual pre-training has demonstrated significant potential in enhancing model performance, particularly in domain-specific scenarios. The most common approach for packing data b…

cs.CV2025

Explainable Synthetic Image Detection through Diffusion Timestep Ensembling

Yixin Wu, Feiran Zhang, Tianyuan Shi +7

Recent advances in diffusion models have enabled the creation of deceptively real images, posing significant security risks when misused. In this study, we empirically show that di…