Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
SGuard-v1: Safety Guardrail for Large Language Models
JoonHo Lee, HyeonMin Cho, Jaewoong Yun +3
We present SGuard-v1, a lightweight safety guardrail for Large Language Models (LLMs), which comprises two specialized models to detect harmful content and screen adversarial promp…
cs.CL2025
Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training Corpora
JoonHo Lee, JuYoun Son, Juree Seok +2
Inconsistent annotations in training corpora, particularly within preference learning datasets, pose challenges in developing advanced language models. These inconsistencies often…