3 papers
cs.CL2026
PaT: Planning-after-Trial for Efficient Test-Time Code Generation
Youngsik Yoon, Sungjae Lee, Seockbean Song +3
Beyond training-time optimization, scaling test-time computation has emerged as a key paradigm to extend the reasoning capabilities of Large Language Models (LLMs). However, most e…
cs.LG2026
Rising Multi-Armed Bandits with Known Horizons
Seockbean Song, Chenyu Gan, Youngsik Yoon +3
The Rising Multi-Armed Bandit (RMAB) framework models environments where expected rewards of arms increase with plays, which models practical scenarios where performance of each op…
cs.LG2024
Combinatorial Rising Bandits
Seockbean Song, Youngsik Yoon, Siwei Wang +2
Combinatorial online learning is a fundamental task for selecting the optimal action (or super arm) as a combination of base arms in sequential interactions with systems providing…