collaborators

7 papers

cs.AI2026

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

Tianjun Pan, Yuan Li, Hongda Wang +8

External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effectiveness depends not only o…

cs.AI2026

Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML Engineering

Shaokang Fu, Yulong Tao, Linbo Jin +7

Long-horizon autonomous research tasks such as machine learning engineering require systems to make interdependent decisions under a limited budget. Existing LLM-based agents typic…

cs.AI2026

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Qiming Shi, Yulong Tao, Linbo Jin +10

Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments…

cs.AI2026

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

Sihang Jiang, Lipeng Ma, Zhonghua Hong +9

Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, failing to accumulate experience…

cs.CL2026

From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks

Qingyu Ren, Tianjun Pan, Xingzhou Chen +1

Large language models have achieved remarkable progress in text generation but still struggle with generative writing tasks. In terms of evaluation, existing benchmarks evaluate wr…

cs.CL2026

OPSDL: On-Policy Self-Distillation for Long-Context Language Models

Xinsen Zhang, Zhenkai Ding, Tianjun Pan +4

Extending the effective context length of large language models (LLMs) remains a central challenge for real-world applications. While recent post-training methods have made progres…