collaborators

7 papers

cs.AI2026

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration

Qifan Zhang, Dongyang Ma, Tianqing Fang +5

Most agents today ``self-evolve'' by following rewards and rules defined by humans. However, this process remains fundamentally dependent on external supervision; without human gui…

cs.AI2026

Exposing Weaknesses of Large Reasoning Models through Graph Algorithm Problems

Qifan Zhang, Jianhao Ruan, Aochuan Chen +4

Large Reasoning Models (LRMs) have advanced rapidly; however, existing benchmarks in mathematics, code, and common-sense reasoning remain limited. They lack long-context evaluation…

cs.CL2026

Incentivizing In-depth Reasoning over Long Contexts with Process Advantage Shaping

Miao Peng, Weizhou Shen, Nuo Chen +3

Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective in enhancing LLMs short-context reasoning, but its performance degrades in long-context scenarios that re…

cs.HC2025

Is AI mingling or bullying me? Exploring User Interactions with a Chatbot in China

Nuo Chen, Pu Yan, Jia Li +1

Since its viral emergence in early 2024, Comment Robert-a Weibo-launched social chatbot-has gained widespread attention on the Chinese Internet for its unsolicited and unpredictabl…

cs.CL2025

Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners

Miao Peng, Nuo Chen, Zongrui Suo +1

Despite significant advancements in Large Language Models (LLMs), developing advanced reasoning capabilities in LLMs remains a key challenge. Process Reward Models (PRMs) have demo…

cs.CL2025

How does Misinformation Affect Large Language Model Behaviors and Preferences?

Miao Peng, Nuo Chen, Jianheng Tang +1

Large Language Models (LLMs) have shown remarkable capabilities in knowledge-intensive tasks, while they remain vulnerable when encountering misinformation. Existing studies have e…