collaborators

6 papers

cs.CL2026

Beyond Cross-Lingual Transfer: Benchmarking Propagation Boundaries in Multilingual LLM Unlearning

Pengyang Shao, Chuanpeng Lu, Wei Qin +5

Large Language Model (LLM) unlearning aims to suppress target knowledge while preserving general capabilities. In multilingual settings, unlearning must additionally propagate with…

cs.LG2026

CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment

Jilong Liu, Yonghui Yang, Pengyang Shao +5

Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes training sensitive to imperfect sup…

cs.LG2026

Controllable Value Alignment in Large Language Models through Neuron-Level Editing

Yonghui Yang, Yihui Wang, Junwei Li +6

Aligning large language models (LLMs) with human values has become increasingly important as their influence on human behavior and decision-making expands. However, existing steeri…

cs.LG2026

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control

Yonghui Yang, Wenjian Tao, Jilong Liu +6

Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment methods focus on uncertainty in alignm…

cs.AI2025

Debate over Mixed-knowledge: A Robust Multi-Agent Reasoning Framework for Incomplete Knowledge Graph Question Answering

Jilong Liu, Pengyang Shao, Wei Qin +3

Knowledge Graph Question Answering (KGQA) aims to improve factual accuracy by leveraging structured knowledge. However, real-world Knowledge Graphs (KGs) are often incomplete, lead…

cs.IR2025

Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation

Yu Wang, Yonghui Yang, Le Wu +3

Recent advances in Large Language Models (LLMs) have opened new avenues for sequential recommendation by enabling natural language reasoning over user behavior sequences. A common…