4 papers
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
Alaa Elsetohy, Sama Hadhoud, Haryo Akbarianto Wibowo +4
Multilingual benchmarks rarely test reasoning over culturally grounded premises: translated datasets keep English-centric scenarios, while culture-first datasets often lack control…
Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming
Sama Hadhoud, Alaa Elsetohy, Frederikus Hudi +3
Large Language Models (LLMs) increasingly succeed on competitive programming problems, yet existing evaluations conflate algorithmic reasoning with code-level implementation. We ar…
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
Haryo Akbarianto Wibowo, Alaa Elsetohy, Qinrong Cui +1
The rapid advancement of Large Language Models (LLMs) has necessitated more robust evaluation methods that go beyond static benchmarks, which are increasingly prone to data saturat…
Khattat: Enhancing Readability and Concept Representation of Semantic Typography
Ahmed Hussein, Alaa Elsetohy, Sama Hadhoud +3
Designing expressive typography that visually conveys a word's meaning while maintaining readability is a complex task, known as semantic typography. It involves selecting an idea,…