collaborators

12 papers

cs.LG2026

Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning

Kishan Panaganti, Zhenwen Liang, Wenhao Yu +2

Recent progress in Large Language Model (LLM) reasoning is increasingly driven by the refinement of post-training loss functions and alignment strategies. However, standard Reinfor…

cs.LG2025

Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning

Zhenwen Liang, Sidi Lu, Wenhao Yu +4

Reinforcement learning has become essential for strengthening the reasoning abilities of large language models, yet current exploration mechanisms remain fundamentally misaligned w…

cs.AI2025

Guided Self-Evolving LLMs with Minimal Human Supervision

Wenhao Yu, Zhenwen Liang, Chengsong Huang +4

AI self-evolution has long been envisioned as a path toward superintelligence, where models autonomously acquire, refine, and internalize knowledge from their own learning experien…

cs.CL2025

Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling

Hyunji Lee, Wenhao Yu, Hongming Zhang +4

Hybrid models that combine state space models (SSMs) with attention mechanisms have shown strong performance by leveraging the efficiency of SSMs and the high recall ability of att…

cs.CL2025

Don't Throw Away Your Pretrained Model

Shangbin Feng, Wenhao Yu, Yike Wang +3

Alignment training has tradeoffs: it helps language models (LMs) gain in reasoning and instruction following but might lose out on skills such as creativity and calibration, where…

cs.CL2025

Retrieval-augmented GUI Agents with Generative Guidelines

Ran Xu, Kaixin Ma, Wenhao Yu +4

GUI agents powered by vision-language models (VLMs) show promise in automating complex digital tasks. However, their effectiveness in real-world applications is often limited by sc…