3 papers
cs.AI2026
CORE: Collaborative Reasoning via Cross Teaching
Kshitij Mishra, Mirat Aubakirov, Martin Takac +2
Large language models exhibit complementary reasoning errors: on the same instance, one model may succeed with a particular decomposition while another fails. We propose Collaborat…
cs.CL2026
SD-E: Semantic Exploration for Reasoning Under Token Budgets
Kshitij Mishra, Nils Lukas, Salem Lahlou
Small language models (SLMs) struggle with complex reasoning because exploration is expensive under tight compute budgets. We introduce Semantic Diversity-Exploration-Exploitation…
cs.LG2025
Tabular and Deep Reinforcement Learning for Gittins Index
Harshit Dhankhar, Kshitij Mishra, Tejas Bodas
In the realm of multi-arm bandit problems, the Gittins index policy is known to be optimal in maximizing the expected total discounted reward obtained from pulling the Markovian ar…