activity
20242026
collaborators

9 papers

cs.LG2026

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

Yunhao Yang, Yuexin Bian, Yunjie Tian +6

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on g…

cs.LG2026

RN-D: Discretized Categorical Actors for On-Policy Reinforcement Learning

Yuexin Bian, Jie Feng, Tao Wang +3

On-policy Reinforcement Learning (RL) remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies,…

cs.LG2026

Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning

Yuan Zhuang, Yuexin Bian, Sihong He +7

Scaling critic capacity is a promising direction for improving off-policy reinforcement learning (RL). However, recent work shows that larger critics are prone to overfitting and i…

cs.CL2026

LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts

Yuan Zhuang, Yi Shen, Yuexin Bian +4

Recent studies have shown that combining parameter-efficient fine-tuning (PEFT) with mixture-of-experts (MoE) is an effective strategy for adapting large language models (LLMs) to…

eess.SY2025

Operator learning for energy-efficient building ventilation control with computational fluid dynamics simulation of a real-world classroom

Yuexin Bian, Oliver Schmidt, Yuanyuan Shi

Energy-efficient ventilation control plays a vital role in reducing building energy consumption while ensuring occupant health and comfort. While Computational Fluid Dynamics (CFD)…

eess.SY2025

DiffOP: Reinforcement Learning of Optimization-Based Control Policies via Implicit Policy Gradients

Yuexin Bian, Jie Feng, Yuanyuan Shi

Real-world control systems require policies that are not only high-performing but also interpretable and robust. A promising direction toward this goal is model-based control, whic…