collaborators

6 papers

cs.CL2025

RCScore: Quantifying Response Consistency in Large Language Models

Dongjun Jang, Youngchae Ahn, Hyopil Shin

Current LLM evaluations often rely on a single instruction template, overlooking models' sensitivity to instruction style-a critical aspect for real-world deployments. We present R…

cs.CL2025

P-CoT: A Pedagogically-motivated Participatory Chain-of-Thought Prompting for Phonological Reasoning in LLMs

Dongjun Jang, Youngchae Ahn, Hyopil Shin

This study explores the potential of phonological reasoning within text-based large language models (LLMs). Utilizing the PhonologyBench benchmark, we assess tasks like rhyme word…

cs.CL2025

KoBALT: Korean Benchmark For Advanced Linguistic Tasks

Hyopil Shin, Sangah Lee, Dongjun Jang +9

We introduce KoBALT (Korean Benchmark for Advanced Linguistic Tasks), a comprehensive linguistically-motivated benchmark comprising 700 multiple-choice questions spanning 24 phenom…

cs.CL2025

MoFE: Mixture of Frozen Experts Architecture

Jean Seo, Jaeyoon Kim, Hyopil Shin

We propose the Mixture of Frozen Experts (MoFE) architecture, which integrates Parameter-efficient Fine-tuning (PEFT) and the Mixture of Experts (MoE) architecture to enhance both…

cs.CL2025

How does a Language-Specific Tokenizer affect LLMs?

Jean Seo, Jaeyoon Kim, SungJoo Byun +1

The necessity of language-specific tokenizers intuitively appears crucial for effective natural language processing, yet empirical analyses on their significance and underlying rea…

cs.CL2024

DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine

Jean Seo, Jongwon Lim, Dongjun Jang +1

We introduce DAHL, a benchmark dataset and automated evaluation system designed to assess hallucination in long-form text generation, specifically within the biomedical domain. Our…