collaborators

6 papers

cs.CL2026

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models

Zongji Yu, Wenshui Luo, Yiliu Sun +4

Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Group Relative Policy Optimizat…

cs.CL2026

TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization

Peiji Li, Linyang Li, Handa Sun +15

Large language models have demonstrated strong reasoning capabilities in complex tasks through tool integration, which is typically framed as a Markov Decision Process and optimize…

cs.CL2025

Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning

Yiliu Sun, Zicheng Zhao, Yang Wei +2

Reinforcement Learning with Verifiable Rewards (RLVR) significantly enhances the reasoning capability of Large Language Models (LLMs). Current RLVR approaches typically conduct tra…

cs.AI2025

CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate

Yiliu Sun, Zicheng Zhao, Sheng Wan +1

Nowadays, single Large Language Model (LLM) struggles with critical issues such as hallucination and inadequate reasoning abilities. To mitigate these issues, Multi-Agent Debate (M…

cs.CL2025

Fast-Slow-Thinking: Complex Task Solving with Large Language Models

Yiliu Sun, Yanfang Zhang, Zicheng Zhao +3

Nowadays, Large Language Models (LLMs) have been gradually employed to solve complex tasks. To face the challenge, task decomposition has become an effective way, which proposes to…

cs.CL2025

Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning

Yanfang Zhang, Yiliu Sun, Yibing Zhan +3

Recently, increasing attention has been focused on improving the ability of Large Language Models (LLMs) to perform complex reasoning. Advanced methods, such as Chain-of-Thought (C…