collaborators

9 papers

cs.CV2026

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Ziyu Ma, Hailang Huang, Shun Zou +5

Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, exis…

cs.CV2026

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration

Ziyu Ma, Shidong Yang, Yuxiang Ji +5

High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when the image is processed as a w…

cs.AI2026

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning

Xucong Wang, Ziyu Ma, Yong Wang +5

Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods of…

cs.LG2026

APPO: Agentic Procedural Policy Optimization

Xucong Wang, Ziyu Ma, Yong Wang +5

Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing metho…

cs.AI2026

Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution

Xucong Wang, Ziyu Ma, Shidong Yang +4

Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static tra…

cs.CL2026

Learning Agentic Policy from Action Guidance

Yuxiang Ji, Zengbin Wang, Yong Wang +6

Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training signals emerge only within its…