3 papers
cs.CL2026
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
Jiashu Yao, Heyan Huang, Zeming Liu +1
To overcome the sparse reward challenge in reinforcement learning (RL) for agents based on large language models (LLMs), we propose Mutual Information Self-Evaluation (MISE), an RL…
cs.CL2026
MemEvolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation
Zihao Cheng, Zeming Liu, Yingyu Shan +7
While large language model--powered agents can self-evolve by accumulating experience or by dynamically creating new assets (i.e., tools or expert agents), existing frameworks typi…
cs.CL2024
ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks
Jiashu Yao, Heyan Huang, Zeming Liu +4
Following formatting instructions to generate well-structured content is a fundamental yet often unmet capability for large language models (LLMs). To study this capability, which…