6 papers
Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection
Yihui Wang, Yonghui Yang, Jilong Liu +3
Deepfake detection suffers from poor generalization across forgery methods, as existing models tend to rely on spurious method-specific shortcuts that fail to transfer to unseen ma…
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
Jilong Liu, Yonghui Yang, Pengyang Shao +5
Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes training sensitive to imperfect sup…
Controllable Value Alignment in Large Language Models through Neuron-Level Editing
Yonghui Yang, Yihui Wang, Junwei Li +6
Aligning large language models (LLMs) with human values has become increasingly important as their influence on human behavior and decision-making expands. However, existing steeri…
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
Yonghui Yang, Wenjian Tao, Jilong Liu +6
Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment methods focus on uncertainty in alignm…
Debate over Mixed-knowledge: A Robust Multi-Agent Reasoning Framework for Incomplete Knowledge Graph Question Answering
Jilong Liu, Pengyang Shao, Wei Qin +3
Knowledge Graph Question Answering (KGQA) aims to improve factual accuracy by leveraging structured knowledge. However, real-world Knowledge Graphs (KGs) are often incomplete, lead…
InterMind: Doctor-Patient-Family Interactive Depression Assessment Empowered by Large Language Models
Zhiyuan Zhou, Jilong Liu, Sanwang Wang +3
Depression poses significant challenges to patients and healthcare organizations, necessitating efficient assessment methods. Existing paradigms typically focus on a patient-doctor…