collaborators

16 papers

cs.CL2026

QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable Rewards

Rongzhi Zhang, Rui Feng, Zhihan Zhang +8

Rubric-based RL is a promising route for extending reinforcement learning beyond verifiable rewards, yet existing methods optimize rubrics while treating the query distribution as…

cs.CL2026

Instant Personalized Large Language Model Adaptation via Hypernetwork

Zhaoxuan Tan, Zixuan Zhang, Haoyang Wen +8

Personalized large language models (LLMs) tailor content to individual preferences using user profiles or histories. However, existing parameter-efficient fine-tuning (PEFT) method…

stat.ML2026

How Accurately Can a Gaussian Approximate Stochastic Approximation Iterates?

Shaan Ul Haque, Zedong Wang, Zixuan Zhang +1

Stochastic approximation (SA) is a method for finding the root of an operator perturbed by noise. The focus of this paper is studying the distribution of SA iterates in finite time…

cs.MA2026

Optimistic ε-Greedy Exploration for Cooperative Multi-Agent Reinforcement Learning

Ruoning Zhang, Siying Wang, Wenyu Chen +5

The Centralized Training with Decentralized Execution (CTDE) paradigm is widely used in cooperative multi-agent reinforcement learning. However, conventional methods based on CTDE…

cs.LG2026

Diffusion Model for Manifold Data: Score Decomposition, Curvature, and Statistical Complexity

Zixuan Zhang, Kaixuan Huang, Tuo Zhao +2

Diffusion models have become a leading framework in generative modeling, yet their theoretical understanding -- especially for high-dimensional data concentrated on low-dimensional…

cs.LG2026

BackPlay: Head-Only Look-Back Self-Correction for Diffusion Language Models

Liming Liu, Binxuan Huang, Zixuan Zhang +3

Diffusion Language Models (DLMs) decode multiple tokens in parallel, but aggressive multi-token decoding amplifies cross-token dependency errors and can sharply degrade generation…