4 papers
Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets
Mykola Khandoga, Yevhen Kostiuk, Anton Polishko +4
Every information ecosystem produces beliefs that shape strategic decisions. Both human analysts and AI systems inherit the blind spots of their information sources. We show that L…
AlignTune: Modular Toolkit for Post-Training Alignment of Large Language Models
R E Zera Marveen Lyngkhoi, Chirag Chawla, Pratinav Seth +5
Post-training alignment is central to deploying large language models (LLMs), yet practical workflows remain split across backend-specific tools and ad-hoc glue code, making experi…
Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization
Mykola Khandoga, Rui Yuan, Vinay Kumar Sankarapu
Policy gradient methods for language model reasoning, such as GRPO and DAPO, assign uniform credit to all generated tokens - the filler phrase "Let me think" receives the same grad…
From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages
Artur Kiulian, Anton Polishko, Mykola Khandoga +10
In this paper, we propose a model-agnostic cost-effective approach to developing bilingual base large language models (LLMs) to support English and any target language. The method…