2 papers
cs.CL2026
CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward
Md Amirul Islam, Sumiran Thakur, Huancheng Chen +3
We present CacheRL, a system for training small agent foundation models that achieves 92 percent process accuracy on multi-step tool-calling tasks, approaching GPT-5's 94 percent w…
cs.LG2025
SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models
Gyuhak Kim, Sumiran Singh Thakur, Su Min Park +2
Supervised fine-tuning (SFT) has become an essential step in tailoring large language models (LLMs) to align with human expectations and specific downstream tasks. However, existin…