activity
20232026
most citedAutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

11 citations · 24 across the 33 of their papers we have counts for

collaborators
Showing 2026Show all

14 papers · 1 filter

cs.LG2026

Pin Once, Swap Light: Subspace-Aligned Centroid-Residual Training for Efficient Ultra-LoRA Serving

Xiang Li, Pengcheng Wang, Huazheng Wang +1

Modern multi-tenant Low-Rank Adapters (LoRAs) serving systems concurrently host tens to hundreds of LoRA adapters. Though powerful, this introduces a critical system dilemma betwee…

cs.AI2026

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Qinsi Wang, Jing Shi, Huazheng Wang +8

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, i…

cs.LG2026

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification

Haoyang Hong, Zichen Wang, Quanquan Gu +1

We study KL-regularized contextual bandits and episodic reinforcement learning (RL) under general function approximation with model misspecification. Existing guarantees rely on re…

cs.AI2026

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?

Jiale Liu, Huajun Xi, Shaokun Zhang +6

Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution i…

cs.CL2026

EvoPool: Evolutionary Programmatic Annotation for Label-Efficient Specialized Supervision

Tianyi Xu, Yaolun Zhang, Xuan Ouyang +1

Large language models excel at general tasks but underperform smaller supervised models in specialized, high-stakes domains where training labels are costly. We address this regime…

cs.AI2026

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

Yifan Zeng, Yiran Wu, Yaolun Zhang +4

Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that…