collaborators

7 papers

cs.AI2026

ENVS: Environment-Native Verified Search for Long-Horizon GUI Agents

Yincheng Zhou, Athena Zhuoming Zhong, Shijie Zhang +3

As multimodal agents move from interface understanding to real software control, successful trajectory discovery in live desktop environments becomes a key challenge. GUI tasks req…

cs.AI2026

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning

Shijie Zhang, Zheng Xiao, Shiyu Liu +7

Online reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning abilities of large language models, but most methods still…

cs.AI2026

Know What You Know: Metacognitive Entropy Calibration for Verifiable RL Reasoning

Qiannian Zhao, Chen Yang, Jinhao Jing +5

Large reasoning models (LRMs) have emerged as a powerful paradigm for solving complex real-world tasks. In practice, these models are predominantly trained via Reinforcement Learni…

cs.CL2026

Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling

Shiqi Yan, Yubo Chen, Ruiqi Zhou +8

The reasoning process of Large Language Models (LLMs) is often plagued by hallucinations and missing facts in question-answering tasks. A promising solution is to ground LLMs' answ…

cs.LG2026

Answer First, Reason Later: Aligning Search Relevance via Mode-Balanced Reinforcement Learning

Shijie Zhang, Xiang Guo, Rujun Guo +4

Building a search relevance model that achieves both low latency and high performance is a long-standing challenge in the search industry. To satisfy the millisecond-level response…

cs.LG2026

ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization

Shijie Zhang, Kevin Zhang, Zheyuan Gu +5

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an important paradigm for unlocking reasoning capabilities in large language models, exemplified by the success…