collaborators

7 papers

cs.CL2026

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…

cs.CL2026

Do Language Models Track Entities Across State Changes?

Zilu Tang, Qiao Zhao, Gabriel Franco +4

Entity tracking (ET), the ability to keep track of states, is a fundamental skill that underlies complex reasoning. An increasing amount of work investigates how transformer langua…

cs.CL2026

mR3: Multilingual Rubric-Agnostic Reward Reasoning Models

David Anugraha, Shou-Yi Hung, Zilu Tang +3

Evaluation using Large Language Model (LLM) judges has been widely adopted in English and shown to be effective for automatic evaluation. However, their performance does not genera…

cs.CL2025

Is Active Persona Inference Necessary for Aligning Small Models to Personal Preferences?

Zilu Tang, Afra Feyza Akyürek, Ekin Akyürek +1

A prominent issue in aligning language models (LMs) to personalized preferences is underspecification -- the lack of information from users about their preferences. A popular trend…

cs.CL2025

R3: Robust Rubric-Agnostic Reward Models

David Anugraha, Zilu Tang, Lester James V. Miranda +5

Reward models are essential for aligning language model outputs with human preferences, yet existing approaches often lack both controllability and interpretability. These models a…

cs.CL2025

A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information

Lucky Susanto, Musa Wijanarko, Prasetia Pratama +6

Polarization is defined as divisive opinions held by two or more groups on substantive issues. As the world's third-largest democracy, Indonesia faces growing concerns about the in…