2 citations · 6 across the 23 of their papers we have counts for
23 papers
Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations
Adril Putra Merin, David Anugraha, Ayu Purwarianti +1
Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmarks evaluate agents within a s…
PingPong: A Natural Benchmark for Multi-Turn Code-Switching Dialogues
Mohammad Rifqi Farhansyah, Hanif Muhammad Zhafran, Farid Adilazuarda +6
Code-switching is a widespread practice among the world's multilingual majority, yet few benchmarks accurately reflect its complexity in everyday communication. We present PingPong…
Beyond Transfer Accuracy: Mechanism-Guided Controlled Adaptation for Low-Resource Languages
Khumaisa Nur'aini, Ayu Purwarianti, Alham Fikri Aji +1
Existing circuit discovery methods rely on templated tasks with clean counterfactuals, limiting their use on diverse natural text. We adapt Contextual Decomposition for Transformer…
Moment and Highlight Detection via MLLM Frame Segmentation
I Putu Andika Bagas Jiwanta, Ayu Purwarianti
Detecting video moments and highlights from natural-language queries have been unified by transformer-based methods. Other works use generative Multimodal LLM (MLLM) to predict mom…
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
IndoPref: A Multi-Domain Pairwise Preference Dataset for Indonesian
Vanessa Rebecca Wiyono, David Anugraha, Ayu Purwarianti +1
Over 200 million people speak Indonesian, yet the language remains significantly underrepresented in preference-based research for large language models (LLMs). Most existing multi…