7 papers
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
Houcheng Jiang, Zetong Zhao, Junfeng Fang +5
Safety-aligned large language models (LLMs) remain vulnerable to backdoor attacks. Recent model editing-based approaches enable efficient backdoor injection by directly modifying a…
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
Ruipeng Wang, Yuxin Chen, Yukai Wang +9
Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployment…
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
Junfeng Fang, Zijun Yao, Ruipeng Wang +3
The development of large language models (LLMs) has entered in a experience-driven era, flagged by the emergence of environment feedback-driven learning via reinforcement learning…
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
Junfeng Fang, Yukai Wang, Ruipeng Wang +5
The rapid advancement of multi-modal large reasoning models (MLRMs) -- enhanced versions of multimodal language models (MLLMs) equipped with reasoning capabilities -- has revolutio…
ACE: Concept Editing in Diffusion Models without Performance Degradation
Ruipeng Wang, Junfeng Fang, Jiaqi Li +4
Diffusion-based text-to-image models have demonstrated remarkable capabilities in generating realistic images, but they raise societal and ethical concerns, such as the creation of…
MixDec Sampling: A Soft Link-based Sampling Method of Graph Neural Network for Recommendation
Xiangjin Xie, Yuxin Chen, Ruipeng Wang +8
Graph neural networks have been widely used in recent recommender systems, where negative sampling plays an important role. Existing negative sampling methods restrict the relation…