9 citations · 14 across the 13 of their papers we have counts for
14 papers
Macaron: Controlled, Human-Written Benchmark for Multilingual and Multicultural Reasoning via Template-Filling
Alaa Elsetohy, Sama Hadhoud, Haryo Akbarianto Wibowo +4
Multilingual benchmarks rarely test reasoning over culturally grounded premises: translated datasets keep English-centric scenarios, while culture-first datasets often lack control…
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
Haryo Akbarianto Wibowo, Alaa Elsetohy, Qinrong Cui +1
The rapid advancement of Large Language Models (LLMs) has necessitated more robust evaluation methods that go beyond static benchmarks, which are increasingly prone to data saturat…
Unveiling the Influence of Amplifying Language-Specific Neurons
Inaya Rahmanisa, Lyzander Marciano Andrylie, Mahardika Krisna Ihsani +3
Language-specific neurons in LLMs that strongly correlate with individual languages have been shown to influence model behavior by deactivating them. However, their role in amplifi…
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
Lyzander Marciano Andrylie, Inaya Rahmanisa, Mahardika Krisna Ihsani +3
Understanding the multilingual mechanisms of large language models (LLMs) provides insight into how they process different languages, yet this remains challenging. Existing studies…
IteRABRe: Iterative Recovery-Aided Block Reduction
Haryo Akbarianto Wibowo, Haiyue Song, Hideki Tanaka +3
Large Language Models (LLMs) have grown increasingly expensive to deploy, driving the need for effective model compression techniques. While block pruning offers a straightforward…
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan +48
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate th…