12 papers
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
Adril Putra Merin, David Anugraha, Ayu Purwarianti +1
Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmarks evaluate agents within a s…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
Khumaisa Nur'aini, Ayu Purwarianti, Alham Fikri Aji +1
Existing circuit discovery methods rely on templated tasks with clean counterfactuals, limiting their use on diverse natural text. We adapt Contextual Decomposition for Transformer…
PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues
Mohammad Rifqi Farhansyah, Hanif Muhammad Zhafran, Farid Adilazuarda +6
Code-switching is a widespread practice among the world's multilingual majority, yet few benchmarks accurately reflect its complexity in everyday communication. We present PingPong…
Moment and Highlight Detection via MLLM Frame Segmentation
I Putu Andika Bagas Jiwanta, Ayu Purwarianti
Detecting video moments and highlights from natural-language queries have been unified by transformer-based methods. Other works use generative Multimodal LLM (MLLM) to predict mom…
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
Vanessa Rebecca Wiyono, David Anugraha, Ayu Purwarianti +1
Over 200 million people speak Indonesian, yet the language remains significantly underrepresented in preference-based research for large language models (LLMs). Most existing multi…