activity
20242026
most citedAnchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.AI2026

Orchestrating Intelligence: Confidence-Aware Routing for Efficient Multi-Agent Collaboration across Multi-Scale Models

Jingbo Wang, Sendong Zhao, Jiatong Liu +4

While multi-agent systems (MAS) have demonstrated superior performance over single-agent approaches in complex reasoning tasks, they often suffer from significant computational ine…

cs.CL2025

Uncovering the Role of Initial Saliency in U-Shaped Attention Bias: Scaling Initial Token Weight for Enhanced Long-Text Processing

Zewen Qiang, Sendong Zhao, Haochun Wang +2

Large language models (LLMs) have demonstrated strong performance on a variety of natural language processing (NLP) tasks. However, they often struggle with long-text sequences due…

cs.CL2025

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security

Yanrui Du, Fenglei Fan, Sendong Zhao +3

As Large Language Models (LLMs) increasingly permeate human life, their security has emerged as a critical concern, particularly their ability to maintain harmless responses to mal…

cs.CL20251 cited

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint

Yanrui Du, Fenglei Fan, Sendong Zhao +6

Instruction Fine-Tuning (IFT) has been widely adopted as an effective post-training strategy to enhance various abilities of Large Language Models (LLMs). However, prior studies ha…

cs.MA2025

Beyond Frameworks: Unpacking Collaboration Strategies in Multi-Agent Systems

Haochun Wang, Sendong Zhao, Jingbo Wang +3

Multi-agent collaboration has emerged as a pivotal paradigm for addressing complex, distributed tasks in large language model (LLM)-driven applications. While prior research has fo…

cs.CL2024

Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning

Yanrui Du, Sendong Zhao, Jiawei Cao +6

Instruction fine-tuning has emerged as a critical technique for customizing Large Language Models (LLMs) to specific applications. However, recent studies have highlighted signific…