4 papers
The Role of Feedback Alignment in Self-Distillation
Semih Kara, OÄuzhan Ersoy
Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains the model to retain this impr…
Congestion Reduction in EV Charger Placement Using Traffic Equilibrium Models
Semih Kara, Yasin Sonmez, Can Kizilkale +3
Growing EV adoption can worsen traffic conditions if chargers are sited without regard to their impact on congestion. We study how to strategically place EV chargers to reduce cong…
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing
Jeffrey Amico, Gabriel Passamani Andrade, John Donaghy +12
Post-training language models (LMs) with reinforcement learning (RL) can enhance their complex reasoning capabilities without supervised fine-tuning, as demonstrated by DeepSeek-R1…
Aggregate Fictitious Play for Learning in Anonymous Polymatrix Games (Extended Version)
Semih Kara, Tamer BaÅar
Fictitious play (FP) is a well-studied algorithm that enables agents to learn Nash equilibrium in games with certain reward structures. However, when agents have no prior knowledge…