10 papers
TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics
Heechan Lee, Jeonggyu Kang, Junho Myung +3
Group conversations are fundamental to human collaboration, yet standard large language models (LLMs) still struggle with the complexities of multi-party interaction. This challeng…
Evalet: Evaluating Large Language Models through Functional Fragmentation
Tae Soo Kim, Heechan Lee, Yoonjoo Lee +2
Practitioners increasingly rely on Large Language Models (LLMs) to evaluate generative AI outputs through "LLM-as-a-Judge" approaches. However, these methods produce holistic score…
SLALOM: Simulation Lifecycle Analysis via Longitudinal Observation Metrics for Social Simulation
Juhoon Lee, Joseph Seering
Large Language Model (LLM) agents offer a potentially-transformative path forward for generative social science but face a critical crisis of validity. Current simulation evaluatio…
Fostering Collective Discourse: A Distributed Role-Based Approach to Online News Commenting
Yoojin Hong, Yersultan Doszhan, Joseph Seering
Current news commenting systems are designed based on implicitly individualistic assumptions, where discussion is the result of a series of disconnected opinions. This often result…
Botender: Supporting Communities in Collaboratively Designing AI Agents through Case-Based Provocations
Tzu-Sheng Kuo, Sophia Liu, Quan Ze Chen +4
AI agents, or bots, serve important roles in online communities. However, they are often designed by outsiders or a few tech-savvy members, leading to bots that may not align with…
AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI
Chae-Gyun Lim, Seung-Ho Han, EunYoung Byun +51
The rapid evolution of generative AI necessitates robust safety evaluations. However, current safety datasets are predominantly English-centric, failing to capture specific risks i…