collaborators

12 papers

cs.LG2026

LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

Zhixin Zhang, Xinke Jiang, Zhibang Yang +5

Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is r…

cs.AI2026

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning

Wentao Zhang, Haoyu Zhang, Xinke Jiang +7

Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contributions of individual reasoning paths. Man…

cs.CL2026

KbSD: Knowledge Boundary aware Self-Distillation for Behavioral Calibration in Agentic Search

Tao Feng, Xinke Jiang, Chao Wu

Agentic search equips large language models with dynamic retrieval abilities, but existing reinforcement learning methods remain limited by reward sparsity in knowledge boundary ca…

cs.CL2026

ProMed: Shapley Information Gain Guided Reinforcement Learning for Proactive Medical LLMs

Hongxin Ding, Baixiang Huang, Yue Fang +8

Interactive medical questioning is essential in clinical consultations, where physicians must actively gather necessary patient information. Yet existing medical Large Language Mod…

cs.AI2026

StackPlanner: A Centralized Hierarchical Multi-Agent System with Task-Experience Memory Management

Ruizhe Zhang, Xinke Jiang, Zhibang Yang +12

Multi-agent systems based on large language models, particularly centralized architectures, have recently shown strong potential for complex and knowledge-intensive tasks. However,…

cs.AI2026

ScaffoldAgent: Utility-Guided Dynamic Outline Optimization for Open-Ended Deep Research

Zhibang Yang, Xinke Jiang, Yuzhen Xiao +9

Open-ended deep research (OEDR) requires systems to acquire knowledge through multi-round retrieval and generate coherent long-form reports. The outline plays a central role as a s…