4 papers
OLA: Output Language Alignment in Code-Switched LLM Interactions
Juhyun Oh, Haneul Yoo, Faiz Ghifari Haznitrama +1
Code-switching, alternating between languages within a conversation, is natural for multilingual users, yet poses fundamental challenges for large language models (LLMs). When a us…
Culture is Everywhere: A Call for Intentionally Cultural Evaluation
Juhyun Oh, Inha Cha, Michael Saxon +3
The prevailing ``trivia-centered paradigm'' for evaluating the cultural alignment of large language models (LLMs) is increasingly inadequate as these models become more advanced an…
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
Jisu Shin, Juhyun Oh, Eunsu Kim +2
Ensuring persona fidelity in large language models (LLMs) is essential for maintaining coherent and engaging human-AI interactions. However, LLMs often exhibit Out-of-Character (OO…
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
Juhyun Oh, Eunsu Kim, Alice Oh
Real-world planning problems require constant adaptation to changing requirements and balancing of competing constraints. However, current benchmarks for evaluating LLMs' planning…