3 papers
cs.CL2025
SGuard-v1: Safety Guardrail for Large Language Models
JoonHo Lee, HyeonMin Cho, Jaewoong Yun +3
We present SGuard-v1, a lightweight safety guardrail for Large Language Models (LLMs), which comprises two specialized models to detect harmful content and screen adversarial promp…
cs.CL2025
Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training Corpora
JoonHo Lee, JuYoun Son, Juree Seok +2
Inconsistent annotations in training corpora, particularly within preference learning datasets, pose challenges in developing advanced language models. These inconsistencies often…
cs.CL2025
Improving Instruction Following in Language Models through Proxy-Based Uncertainty Estimation
JoonHo Lee, Jae Oh Woo, Juree Seok +11
Assessing response quality to instructions in language models is vital but challenging due to the complexity of human language across different contexts. This complexity often resu…