14 papers
FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts
Sizhe Tang, Guangyu Jiang, Yu Li +3
Large Language Models (LLMs) perform strongly on well-specified reasoning tasks with a feasible answer. However, problems encountered in the open world can become ill-posed due to…
Localizing Credit at the Divergence: Path-Conditioned Self-Distillation for LLM Reasoning
Yu Li, Shu Hong, Tian Lan
Reinforcement learning from verifiable rewards assigns a single scalar to each rollout, leaving token-level credit assignment underspecified in long reasoning traces. On-policy sel…
LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition
Yanyu Chen, Jiyue Jiang, Dianzhi Yu +8
The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous rewards offers a solution, m…
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
Rongqian Chen, Yu Li, Zeyu Fang +3
Computer-Use Agents (CUAs) leverage large language models to execute GUI operations on desktop environments, yet they generate actions without evaluating action quality, leading to…
AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery
Yu Li, Chenyang Shao, Xinyang Liu +13
Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creat…
Voting with the Graph: Stable RLAIF via Topological Consistency Maximization
Boyin Liu, Zhuo Zhang, Sen Huang +8
Reinforcement Learning from AI Feedback (RLAIF) relies on LLM judges as preference measurement instruments, yet these instruments are fundamentally limited by random measurement er…