4 papers
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
Junhyuk Choi, Sohhyung Park, Chanhee Cho +2
While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether…
Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
Junhyuk Choi, Hyeonchu Park, Haemin Lee +3
Recent advances in Large Language Models (LLMs) have generated significant interest in their capacity to simulate human-like behaviors, yet most studies rely on fictional personas…
PHISH in MESH: Korean Adversarial Phonetic Substitution and Phonetic-Semantic Feature Integration Defense
Byungjun Kim, Minju Kim, Hyeonchu Park +1
As malicious users increasingly employ phonetic substitution to evade hate speech detection, researchers have investigated such strategies. However, two key challenges remain. Firs…
DART: An AIGT Detector using AMR of Rephrased Text
Hyeonchu Park, Byungjun Kim, Bugeun Kim
As large language models (LLMs) generate more human-like texts, concerns about the side effects of AI-generated texts (AIGT) have grown. So, researchers have developed methods for…