2 papers
cs.LG2026
TVCACHE: A Stateful Tool-Value Cache for Post-Training LLM Agents
Abhishek Vijaya Kumar, Bhaskar Kataria, Byungsoo Oh +2
In RL post-training of LLM agents, calls to external tools take several seconds or even minutes, leaving allocated GPUs idle and inflating post-training time and cost. While many t…
cs.DC2025
FlashMoE: Fast Distributed MoE in a Single Kernel
Osayamen Jonathan Aimuyo, Byungsoo Oh, Rachee Singh
The computational sparsity of Mixture-of-Experts (MoE) models enables sub-linear growth in compute cost as model size increases, thus offering a scalable path to training massive n…