42 citations · 110 across the 36 of their papers we have counts for
39 papers
DriveThru: a Document Extraction Platform and Benchmark Datasets for Indonesian Local Language Archives
Mohammad Rifqi Farhansyah, Muhammad Zuhdi Fikri Johari, Afinzaki Amiral +3
Indonesia is one of the most diverse countries linguistically. However, despite this linguistic diversity, Indonesian languages remain underrepresented in Natural Language Processi…
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
David Anugraha, Garry Kuwanto, Lucky Susanto +2
We present MetaMetrics-MT, an innovative metric designed to evaluate machine translation (MT) tasks by aligning closely with human preferences through Bayesian optimization with Ga…
Linguistics Theory Meets LLM: Code-Switched Text Generation via Equivalence Constrained Large Language Models
Garry Kuwanto, Chaitanya Agarwal, Genta Indra Winata +1
Code-switching, the phenomenon of alternating between two or more languages in a single conversation, presents unique challenges for Natural Language Processing (NLP). Most existin…
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan +48
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate th…
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
Genta Indra Winata, David Anugraha, Lucky Susanto +2
Understanding the quality of a performance evaluation metric is crucial for ensuring that model outputs align with human preferences. However, it remains unclear how well each metr…
Generating Faithful and Salient Text from Multimodal Data
Tahsina Hashem, Weiqing Wang, Derry Tanti Wijaya +2
While large multimodal models (LMMs) have obtained strong performance on many multimodal tasks, they may still hallucinate while generating text. Their performance on detecting sal…