2 papers
cs.AI2026
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations
Lirong Che, Yuzhe yang, Peiwen lin +3
Agent harness evolution improves frozen language-model agents by modifying the executable structures around them. We study this paradigm as a form of sample-efficient fast adaptati…
cs.LG2025
An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines
Jianhai Su, Jinzhu Luo, Qi Zhang
We take the novel perspective of incorporating offline RL algorithms as subroutines of tabula rasa online RL. This is feasible because an online learning agent can repurpose its hi…