activity
20212026
most citedChatGPT Evaluation on Sentence Level Relations: A Focus on Temporal, Causal, and Discourse Relations

30 citations · 108 across the 57 of their papers we have counts for

collaborators

58 papers

cs.CL2026

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

Rui Wang, Hongru Wang, Yi Chen +4

On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, p…

cs.CL2026

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

Xinyu Geng, Xuanhua He, Sixiang Chen +7

Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward rei…

cs.AI2026

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

Tianshi Zheng, Rui Wang, Xiyun Li +5

Frontier scientific reasoning is rapidly emerging as a key foundation for advancing AI agents in automated scientific discovery. Deep research agents offer a promising approach to…

cs.CL2026

PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments

Jiaxin Bai, Yue Guo, Yifei Dong +13

World models for interactive text agents must typically be learned from observation-action trajectories alone. Specifically, the environment returns text observations after each ac…

cs.LG2026

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding

Wenkai Wang, Xiyun Li, Hongcan Guo +5

Graphical User Interface (GUI) grounding requires mapping natural language instructions to precise pixel coordinates. However, due to visually homogeneous elements and dense layout…

cs.AI2026

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration

Qifan Zhang, Dongyang Ma, Tianqing Fang +5

Most agents today ``self-evolve'' by following rewards and rules defined by humans. However, this process remains fundamentally dependent on external supervision; without human gui…