From the 1 of 4 linked papers with an AI index.
4 papers
Reward-Free Evolving Agents via Pairwise Validator
Minghao Liu, Yu Wang, Jiayun Wang +1
The paper introduces a reward‑free approach for self‑evolving agents by using a frozen large language model as a pairwise validator that decides which of two agent versions is bett…
CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward
Md Amirul Islam, Sumiran Thakur, Huancheng Chen +3
We present CacheRL, a system for training small agent foundation models that achieves 92 percent process accuracy on multi-step tool-calling tasks, approaching GPT-5's 94 percent w…
Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory
Zhenting Wang, Huancheng Chen, Jiayun Wang +1
Large language model (LLM) agents are fundamentally bottlenecked by finite context windows on long-horizon tasks. As trajectories grow, retaining tool outputs and intermediate reas…
Multi-Modal Self-Supervised Learning for Surgical Feedback Effectiveness Assessment
Arushi Gupta, Rafal Kocielnik, Jiayun Wang +5
During surgical training, real-time feedback from trainers to trainees is important for preventing errors and enhancing long-term skill acquisition. Accurately predicting the effec…