13 papers
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
Yihui Wang, Yonghui Yang, Jilong Liu +3
Deepfake detection suffers from poor generalization across forgery methods, as existing models tend to rely on spurious method-specific shortcuts that fail to transfer to unseen ma…
Controllable Value Alignment in Large Language Models through Neuron-Level Editing
Yonghui Yang, Yihui Wang, Junwei Li +6
Aligning large language models (LLMs) with human values has become increasingly important as their influence on human behavior and decision-making expands. However, existing steeri…
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
Jilong Liu, Yonghui Yang, Pengyang Shao +5
Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes training sensitive to imperfect sup…
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
Yonghui Yang, Wenjian Tao, Jilong Liu +6
Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment methods focus on uncertainty in alignm…
A Survey on Generative Recommendation: Data, Model, and Tasks
Min Hou, Le Wu, Yuxin Liao +6
Recommender systems serve as foundational infrastructure in modern information ecosystems, helping users navigate digital content and discover items aligned with their preferences.…
Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
Yu Wang, Yonghui Yang, Le Wu +3
Recent advances in Large Language Models (LLMs) have opened new avenues for sequential recommendation by enabling natural language reasoning over user behavior sequences. A common…