3 papers
cs.CL2025
RCScore: Quantifying Response Consistency in Large Language Models
Dongjun Jang, Youngchae Ahn, Hyopil Shin
Current LLM evaluations often rely on a single instruction template, overlooking models' sensitivity to instruction style-a critical aspect for real-world deployments. We present R…
cs.CL2025
P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs
Dongjun Jang, Youngchae Ahn, Hyopil Shin
This study explores the potential of phonological reasoning within text-based large language models (LLMs). Utilizing the PhonologyBench benchmark, we assess tasks like rhyme word…
cs.CL2025
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
Hyopil Shin, Sangah Lee, Dongjun Jang +9
We introduce KoBALT (Korean Benchmark for Advanced Linguistic Tasks), a comprehensive linguistically-motivated benchmark comprising 700 multiple-choice questions spanning 24 phenom…