activity
20242026
collaborators

6 papers

cs.CL2026

Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation

Huije Lee, Jisu Shin, Hoyun Song +2

Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale pre-training corpora. To add…

cs.CL2026

RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity

Jisu Shin, Hoyun Song, Juhyun Oh +4

People often encounter role conflicts -- social dilemmas where the expectations of multiple roles clash and cannot be simultaneously fulfilled. As large language models (LLMs) incr…

cs.CL2025

Language over Content: Tracing Cultural Understanding in Multilingual Large Language Models

Seungho Cho, Changgeon Ko, Eui Jun Hwang +3

Large language models (LLMs) are increasingly used across diverse cultural contexts, making accurate cultural understanding essential. Prior evaluations have mostly focused on outp…

cs.CL2025

A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings

Fitsum Gaim, Hoyun Song, Huije Lee +3

Content moderation research has recently made significant advances, but remains limited in serving the majority of the world's languages due to the lack of resources, leaving milli…

cs.CL2025

Does Rationale Quality Matter? Enhancing Mental Disorder Detection via Selective Reasoning Distillation

Hoyun Song, Huije Lee, Jisu Shin +3

The detection of mental health problems from social media and the interpretation of these results have been extensively explored. Research has shown that incorporating clinical sym…

cs.CL2024

Different Bias Under Different Criteria: Assessing Bias in LLMs with a Fact-Based Approach

Changgeon Ko, Jisu Shin, Hoyun Song +2

Large language models (LLMs) often reflect real-world biases, leading to efforts to mitigate these effects and make the models unbiased. Achieving this goal requires defining clear…