activity
20242026
collaborators

42 papers

cs.LG2026

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

Ting Zhou, Zhenqing Ling, Daoyuan Chen +4

Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute w…

cs.CR2026

Private Direct Preference Optimization for LLM Alignment

Yangfan Jiang, Fei Wei, Ergute Bao +3

Direct preference optimization (DPO) is now a standard method for aligning large language models (LLMs) using human preference data. Each DPO example contains a prompt and a pair o…

cs.CR2026

Hybrid Analysis for Secure MCP Tool Use in LLM Agents

Ping He, Yuexiang Xie, Yaliang Li +1

The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize interactions between LLM agents and exte…

cs.AI2026

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

Haipeng Ding, Yuexiang Xie, Zhewei Wei +2

Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on stati…

cs.AI2026

CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

Yuchen Huang, Xiang Li, Zhenqing Ling +5

Data refinement involves executing multi-step recipes over evolving text states, where both composition and execution order of processing operators determine the outcome. While exi…

cs.LG2026

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

Yanxi Chen, Weijie Shi, Yuexiang Xie +4

This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based A…