39 papers
CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages
Muhammad Dehan Al Kautsar, Salsabila Pranida, Bilal Elbouardi +1
Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping away the dialogic context in w…
From Prediction to Intervention: Personalized Meal-Level Glucose Regulation via an LLM Agent
Mingyu Huang, Weiqing Min, Ying Jin +2
Personalized glucose regulation remains a central yet unresolved challenge in precision nutrition, as postprandial glucose response varies substantially across individuals. Existin…
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection
Muhammad Alif Al Hakim, Alfan Farizki Wicaksono, Fajri Koto
Quantization is widely adopted to reduce the computational cost of large language models (LLMs); however, its implications for fairness and safety, particularly in dynamic quantiza…
Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation
Nurkhan Laiyk, Gerard I. Gállego, Javier Ferrando +1
Function vectors (FVs) are vector representations of tasks extracted from model activations during in-context learning. While prior work has shown that multilingual model represent…
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
Amir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan +1
While automatic metrics drive progress in Machine Translation (MT) and Text Summarization (TS), existing metrics have been developed and validated almost exclusively for English an…
Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems
Muhammad Falensi Azmi, Ikhlasul Akmal Hanif, Vallerie Alexandra Putra +3
Symbolic benchmarks have emerged as a key approach to assess model robustness under minor modifications to STEM-related questions. However, existing symbolic benchmarks mostly rema…