collaborators

7 papers

cs.LG2026

Physics-Guided Policy Optimization with Self-Distillation

Ke Wang, Yuning Wu, Haoran Liu +3

Self-distilled policy optimization (SDPO) has become a popular paradigm for LLM post-training, where a model learns from its own predictions conditioned on privileged information.…

cs.AI2026

CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Models

Yuning Wu, Yingmin Liu, Yang Shu

Large language model (LLM) self-correction -- the ability to detect and fix errors in generated outputs -- remains largely ad hoc, relying on generic prompts such as "please recons…

cs.CL2026

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

Xuanyu Lei, Chenliang Li, Yuning Wu +7

Recent advances in Large Language Models(LLMs) have enabled strong performance in long-form writing, but current training paradigms remain limited: Supervised Fine-Tuning (SFT) rem…

cs.LG2026

Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings

Yuning Wu, Ke Wang, Devin Chen +1

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for post-training reasoning models. However, group-based methods such as Group Relative Po…

cs.CL2026

R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning

Wanlong Liu, Bo Zhang, Chenliang Li +4

While deep reasoning with long chain-of-thought has dramatically improved large language models in verifiable domains like mathematics, its effectiveness for open-ended tasks such…

cs.SD2025

UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities

Xuenan Xu, Jiahao Mei, Zihao Zheng +9

Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where ea…