papers

Publications (11)

cs.CL2025

Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models

Qingyu Ren, Jie Zeng, Qianyu He +5

It is crucial for large language models (LLMs) to follow instructions that involve multiple constraints. However, it is an unexplored area to enhance LLMs' ability to follow soft c…

cs.AI2024

Minor DPO reject penalty to increase training robustness

Shiming Xie, Hong Chen, Fred Yu +3

Learning from human preference is a paradigm used in large-scale language model (LLM) fine-tuning step to better align pretrained LLM to human preference for downstream task. In th…

cs.CL2026

SEIF: Self-Evolving Reinforcement Learning for Instruction Following

Qingyu Ren, Qianyu He, Jiajie Zhu +7

Instruction following is a fundamental capability of large language models (LLMs), yet continuously improving this capability remains challenging. Existing methods typically rely e…

cs.AI2024

Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation

Shiming Xie, Hong Chen, Fred Yu +2

Instruct LLM provide a paradigm used in large scale language model to align LLM to human preference. The paradigm contains supervised fine tuning and reinforce learning from human…

cs.AI2026

Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking

Jinyi Han, Ying Huang, Ying Liao +11

Large Reasoning Models (LRMs) have achieved impressive performance on challenging tasks, yet their deep reasoning often incurs substantial computational costs. To achieve efficient…

cs.CL2026

Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following

Qingyu Ren, Qianyu He, Powei Chang +5

Language models often struggle to follow multi-constraint instructions that are crucial for real-world applications. Existing reinforcement learning (RL) approaches suffer from dep…