Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Jing Liang, Hongyao Tang, Yi Ma +5
Reinforcement Learning (RL) has demonstrated its potential to improve the reasoning ability of Large Language Models (LLMs). One major limitation of most existing Reinforcement Fin…
cs.LG2025
EvoFlow: Evolving Diverse Agentic Workflows On The Fly
Guibin Zhang, Kaijie Chen, Guancheng Wan +5
The past two years have witnessed the evolution of large language model (LLM)-based multi-agent systems from labor-intensive manual design to partial automation (\textit{e.g.}, pro…