4 papers · 1 filter
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
Yejin Lee, Su-Hyeon Kim, Hyundong Jin +3
As language models become increasingly deployed in online environments, toxicity detection and detoxification have received growing attention. Existing studies primarily focus on n…
Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models
Yejin Lee, Yo-Sub Han
Diffusion Language Models (DLMs) provide a promising alternative to autoregressive language models by generating text through iterative denoising and bidirectional refinement. Howe…
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
Yejin Lee, Hyeseon Ahn, Yo-Sub Han
Hate speech remains prevalent in human society and continues to evolve in its forms and expressions. Modern advancements in internet and online anonymity accelerate its rapid sprea…
AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
Yejin Lee, Joonghyuk Hahn, Hyeseon Ahn +1
Implicit hate speech detection is challenging due to its subtlety and reliance on contextual interpretation rather than explicit offensive words. Current approaches rely on contras…