activity
20242026
collaborators

12 papers

cs.LG2026

Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off

Zhaochun Li, Chen Wang, Jionghao Bai +4

The exploration-exploitation (EE) trade-off is a central challenge in reinforcement learning (RL) for large language models (LLMs). With Group Relative Policy Optimization (GRPO),…

cs.AI2025

How Modality Shapes Perception and Reasoning: A Study of Error Propagation in ARC-AGI

Bo Wen, Chen Wang, Erhan Bilal

ARC-AGI and ARC-AGI-2 measure generalization-through-composition on small color-quantized grids, and their prize competitions make progress on these harder held-out tasks a meaning…

cs.PF2025

A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving

Ferran Agullo, Joan Oliveras, Chen Wang +5

With the rapid adoption of Large Language Models (LLMs), LLM-adapters have become increasingly common, providing lightweight specialization of large-scale models. Serving hundreds…

cs.AI2025

Voice-based AI Agents: Filling the Economic Gaps in Digital Health Delivery

Bo Wen, Chen Wang, Qiwei Han +4

The integration of voice-based AI agents in healthcare presents a transformative opportunity to bridge economic and accessibility gaps in digital health delivery. This paper explor…

cs.LG2025

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning

Yanzhi Zhang, Zhaoxi Zhang, Haoxiang Guan +6

Reinforcement learning has emerged as a powerful paradigm for post-training large language models (LLMs) to improve reasoning. Approaches like Reinforcement Learning from Human Fee…

cs.LG2025

EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework

Chen Wang, Lai Wei, Yanzhi Zhang +5

Recent advances in reinforcement learning (RL) have significantly enhanced the reasoning capabilities of large language models (LLMs). Group Relative Policy Optimization (GRPO), a…