activity
20242026
collaborators

26 papers

cs.LG2026

Pin Once, Swap Light: Subspace-Aligned Centroid-Residual Training for Efficient Ultra-LoRA Serving

Xiang Li, Pengcheng Wang, Huazheng Wang +1

Modern multi-tenant Low-Rank Adapters (LoRAs) serving systems concurrently host tens to hundreds of LoRA adapters. Though powerful, this introduces a critical system dilemma betwee…

cs.AI2026

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Qinsi Wang, Jing Shi, Huazheng Wang +8

Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, i…

cs.LG2026

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification

Haoyang Hong, Zichen Wang, Quanquan Gu +1

We study KL-regularized contextual bandits and episodic reinforcement learning (RL) under general function approximation with model misspecification. Existing guarantees rely on re…

cs.AI2026

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?

Jiale Liu, Huajun Xi, Shaokun Zhang +6

Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution i…

cs.LG2026

When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs

Jose Efraim Aguilar Escamilla, Haoyang Hong, Jiawei Li +4

We study reward poisoning attacks in reinforcement learning (RL), where an adversary manipulates rewards within constrained budgets to force the target RL agent to adopt a policy t…

cs.CL2026

Speculative Pipeline Decoding: Higher-Accuracy Drafting with Hidden Latency via Pipeline Parallelism

Yijiong Yu, Huazheng Wang, Shuai Yuan +2

Speculative Decoding (SD) accelerates low-concurrency LLM inference with a draft-then-verify paradigm. Mainstream methods, however, rely on multi-token prediction, which incurs com…