activity
20242026
collaborators

5 papers

cs.CL2026

LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents

Xiaoxuan Peng, Kaiqi Zhang, Xinyu Lu +5

Mastering terminal environments requires language agents capable of multi-step planning, feedback-grounded execution, and dynamic state adaptation. However, training such agents is…

cs.AI2026

Agent Learning via Early Experience

Kai Zhang, Xiangchao Chen, Bo Liu +27

A long-term goal of language agents is to learn and improve through their own experience, ultimately outperforming humans in complex, real-world tasks. However, training agents fro…

cs.LG2026

P^2O: Joint Policy and Prompt Optimization

Xinyu Lu, Kaiqi Zhang, Jinglin Yang +6

Reinforcement Learning with Verifiable Rewards (RLVR) enhances Large Language Model (LLM) reasoning but suffers from advantage collapse on ``hard samples'' where all rollouts fail.…

cs.AI2025

Scaling Agent Learning via Experience Synthesis

Zhaorun Chen, Zhuokai Zhao, Kai Zhang +15

While reinforcement learning (RL) can empower autonomous agents by enabling self-improvement through interaction, its practical adoption remains challenging due to costly rollouts,…

cs.CL2024

ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue

Zhangpu Li, Changhong Zou, Suxue Ma +13

The rocketing prosperity of large language models (LLMs) in recent years has boosted the prevalence of vision-language models (VLMs) in the medical sector. In our online medical co…