5 papers
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
Kevin Wang, Anna Thöni, Benjamin Kempinski +50
Large language models (LLMs) are increasingly deployed as interactive agents, yet their capacity for social and strategic reasoning over extended interaction remains poorly underst…
MindGames Arena Generalization Track: In2AI Solution with Delayed Per-Step Reward Attribution
Aliaksei Korshuk, Alexander Buyantuev, Ilya Makarov
Training language model agents for multi-agent strategic interaction presents a core difficulty: the quality of any action may depend on future events that never materialize, on mo…
ReDAct: Uncertainty-Aware Deferral for LLM Agents
Dzianis Piatrashyn, Nikita Kotelevskii, Kirill Grishchenkov +7
Recently, LLM-based agents have become increasingly popular across many applications, including complex sequential decision-making problems. However, they inherit the tendency of L…
SuS: Strategy-aware Surprise for Intrinsic Exploration
Mark Kashirskiy, Ilya Makarov
We propose Strategy-aware Surprise (SuS), a novel intrinsic motivation framework that uses pre-post prediction mismatch as a novelty signal for exploration in reinforcement learnin…
AraToken: Optimizing Arabic Tokenization with Normalization Pipeline and Language Extension for Qwen3
Mark Kashirskiy, Artiom Lipinski, Ilya Makarov
Tokenization is a critical preprocessing step for large language models (LLMs), directly impacting training efficiency and downstream performance. General-purpose tokenizers traine…