7 papers
Physics-Guided Policy Optimization with Self-Distillation
Ke Wang, Yuning Wu, Haoran Liu +3
Self-distilled policy optimization (SDPO) has become a popular paradigm for LLM post-training, where a model learns from its own predictions conditioned on privileged information.…
CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Models
Yuning Wu, Yingmin Liu, Yang Shu
Large language model (LLM) self-correction -- the ability to detect and fix errors in generated outputs -- remains largely ad hoc, relying on generic prompts such as "please recons…
Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
Xuanyu Lei, Chenliang Li, Yuning Wu +7
Recent advances in Large Language Models(LLMs) have enabled strong performance in long-form writing, but current training paradigms remain limited: Supervised Fine-Tuning (SFT) rem…
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
Yuning Wu, Ke Wang, Devin Chen +1
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for post-training reasoning models. However, group-based methods such as Group Relative Po…
R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning
Wanlong Liu, Bo Zhang, Chenliang Li +4
While deep reasoning with long chain-of-thought has dramatically improved large language models in verifiable domains like mathematics, its effectiveness for open-ended tasks such…
UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
Xuenan Xu, Jiahao Mei, Zihao Zheng +9
Audio generation, including speech, music and sound effects, has advanced rapidly in recent years. These tasks can be divided into two categories: time-aligned (TA) tasks, where ea…