collaborators

14 papers

cs.AI2026

FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts

Sizhe Tang, Guangyu Jiang, Yu Li +3

Large Language Models (LLMs) perform strongly on well-specified reasoning tasks with a feasible answer. However, problems encountered in the open world can become ill-posed due to…

cs.LG2026

Localizing Credit at the Divergence: Path-Conditioned Self-Distillation for LLM Reasoning

Yu Li, Shu Hong, Tian Lan

Reinforcement learning from verifiable rewards assigns a single scalar to each rollout, leaving token-level credit assignment underspecified in long reasoning traces. On-policy sel…

cs.AI2026

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition

Yanyu Chen, Jiyue Jiang, Dianzhi Yu +8

The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous rewards offers a solution, m…

cs.AI2026

IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents

Rongqian Chen, Yu Li, Zeyu Fang +3

Computer-Use Agents (CUAs) leverage large language models to execute GUI operations on desktop environments, yet they generate actions without evaluating action quality, leading to…

cs.CL2026

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery

Yu Li, Chenyang Shao, Xinyang Liu +13

Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creat…

cs.AI2026

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization

Boyin Liu, Zhuo Zhang, Sen Huang +8

Reinforcement Learning from AI Feedback (RLAIF) relies on LLM judges as preference measurement instruments, yet these instruments are fundamentally limited by random measurement er…