15 papers
TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics
Heechan Lee, Jeonggyu Kang, Junho Myung +3
Group conversations are fundamental to human collaboration, yet standard large language models (LLMs) still struggle with the complexities of multi-party interaction. This challeng…
IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering
JungMin Yun, YoungBin Kim
Multi-hop question answering requires complex reasoning across multiple evidence segments, which often overwhelms retrieval-augmented generation systems with lengthy and noisy cont…
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe +57
Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances,…
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
Jiho Jin, Junho Myung, Juhyun Oh +5
As large language models (LLMs) are increasingly deployed to users around the world, they are integrated into everyday tasks across diverse cultural contexts, from drafting persona…
SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
Nedjma Ousidhoum, Junho Myung, Carla Perez-Almendros +27
We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manual…
FINEST: Improving LLM Responses to Sensitive Topics Through Fine-Grained Evaluation
Juhyun Oh, Nayeon Lee, Chani Jung +5
Large Language Models (LLMs) often generate overly cautious and vague responses on sensitive topics, sacrificing helpfulness for safety. Existing evaluation frameworks lack systema…