1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.CL2025
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
cs.CL2025★ 1 cited
Trainable Reference-Based Evaluation Metric for Identifying Quality of English-Gujarati Machine Translation System
Nisheeth Joshi, Pragya Katyayan, Palak Arora
Machine Translation (MT) Evaluation is an integral part of the MT development life cycle. Without analyzing the outputs of MT engines, it is impossible to evaluate the performance…