3 papers
cs.IR2026
E-SENS: Exclusion-Sensitive Penalization for Negative-Constraint Retrieval
Yerang Kim, Jiyoon Myung, Joohyung Han
Retrieval-augmented language models can fail to respect negative constraints when the retriever supplies evidence about concepts the user explicitly excluded. Beyond explicit negat…
cs.CL2025
K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean
Minkyeong Jeon, Hyemin Jeong, Yerang Kim +3
Language detoxification involves removing toxicity from offensive language. While a neutral-toxic paired dataset provides a straightforward approach for training detoxification mod…
cs.CL2025
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
Negar Foroutan, Jakhongir Saydaliev, Ye Eun Kim +1
Language identification (LID) is a critical step in curating multilingual LLM pretraining corpora from web crawls. While many studies on LID model training focus on collecting dive…