4 papers
TritonRL: Training LLMs to Think and Code Triton Without Cheating
Jiin Woo, Shaowei Zhu, Allen Nie +3
The rapid evolution of Large Language Models (LLMs) has driven a growing demand for automated, high-performance system kernels to accelerate machine learning workloads. We introduc…
Federate the Router: Learning Language Model Routers with Sparse and Decentralized Evaluations
Baris Askin, Shivam Patel, Anupam Nayak +4
Large language models (LLMs) are increasingly accessed as remotely hosted services by edge and enterprise clients that cannot run frontier models locally. Since models vary widely…
Sample Complexity of Average-Reward Q-Learning: From Single-agent to Federated Reinforcement Learning
Yuchen Jiao, Jiin Woo, Gen Li +2
Average-reward reinforcement learning offers a principled framework for long-term decision-making by maximizing the mean reward per time step. Although Q-learning is a widely used…
Large Language Model-Enhanced Reinforcement Learning for Diverse and Novel Recommendations
Jiin Woo, Alireza Bagheri Garakani, Tianchen Zhou +2
In recommendation systems, diversity and novelty are essential for capturing varied user preferences and encouraging exploration, yet many systems prioritize click relevance. While…