2 papers
cs.LG2026
LeAct: Learning to Reason from Expert Actions
Ziran Yang, Chengshuai Shi, Raj Ghugare +3
Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision l…
cs.LG2024
Scaling Laws for Imitation Learning in Single-Agent Games
Jens Tuyls, Dhruv Madeka, Kari Torkkola +3
Imitation Learning (IL) is one of the most widely used methods in machine learning. Yet, many works find it is often unable to fully recover the underlying expert behavior, even in…