1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2026
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
Amir Hossein Yari, Fajri Koto
Reinforcement learning has become the primary paradigm for aligning large language models (LLMs) on complex reasoning tasks, with group relative policy optimization (GRPO) widely u…
cs.CL2025★ 1 cited
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension
Amir Hossein Yari, Fajri Koto
Despite the impressive performance of multilingual large language models (mLLMs) in various natural language processing tasks, their ability to understand procedural texts, particu…