5 papers
Distribution-aware Language Neuron Identification in Multilingual Large Language Models
Minjun Kim, Inho Won, Junghun Yuk +3
Multilingual large language models (mLLMs) contain a small fraction of feed-forward neurons that are sensitive to particular languages, commonly termed language-specific neurons. E…
Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text
Minkyung Cho, Jihyo Kim, SeungWoo Song +4
Synthetic data is increasingly used to train large language models (LLMs), yet its security implications remain poorly understood. Prior work on subliminal learning suggests that m…
TELLME: Test-Enhanced Learning for Language Model Enrichment
Minjun Kim, Inho Won, Hyeonseok Lim +6
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consistently been accompanied by challenges, such…
ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts
Dongwon Noh, Donghyeok Koh, Junghun Yuk +4
Prior benchmarks for evaluating the domain-specific knowledge of large language models (LLMs) lack the scalability to handle complex academic tasks. To address this, we introduce \…
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
Hyeonseok Lim, Dongjae Shin, Seohyun Song +5
We propose the VLR-Bench, a visual question answering (VQA) benchmark for evaluating vision language models (VLMs) based on retrieval augmented generation (RAG). Unlike existing ev…