activity
20222026
most citedLoRA Dropout as a Sparsity Regularizer for Overfitting Control

8 citations · 8 across the 13 of their papers we have counts for

collaborators

15 papers

cs.MA2026

Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses

Xinke Jiang, Zhixin Zhang, Zhibang Yang +6

Large language model agents increasingly solve long-horizon tasks through multi-agent harnesses in which a central agent coordinates specialized sub-agents, tools, and environments…

cs.MA2026

AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing

Xinke Jiang, Yue Fang, Zhibang Yang +12

Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs), yet existing RAG systems often struggle with complex, multi-step reasoning that requir…

cs.LG2026

LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation

Zhixin Zhang, Xinke Jiang, Zhibang Yang +5

Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is r…

cs.SE2026

Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes

Yuhao Tan, Zhibang Yang, Fangkai Yang +9

Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly from bug reports remains underconstrained.…

cs.LG2026

ToolAtlas: Learning Once, Reusing Everywhere with Tool-Side Memory

Yue Fang, Zhibang Yang, Fangkai Yang +5

Large language model (LLM) agents increasingly rely on external tools served by shared providers and accessed by heterogeneous downstream agents. Existing approaches improve tool u…

cs.LG2026

The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment

Tianyu Jia, Yue Fang, Hongxin Ding +6

Process reward models (PRMs) enhance the reasoning capabilities of large language models (LLMs) by providing fine-grained feedback, yet training PRMs typically requires expensive s…