3 papers
cs.LG2024
Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations
Swapnaja Achintalwar, Adriana Alvarado Garcia, Ateret Anaby-Tavor +35
Large language models (LLMs) are susceptible to a variety of risks, from non-faithful output to biased and toxic generations. Due to several limiting factors surrounding LLMs (trai…
cs.CL2024
Efficient Models for the Detection of Hate, Abuse and Profanity
Christoph Tillmann, Aashka Trivedi, Bishwaranjan Bhattacharjee
Large Language Models (LLMs) are the cornerstone for many Natural Language Processing (NLP) tasks like sentiment analysis, document classification, named entity recognition, questi…
cs.CL2023
Muted: Multilingual Targeted Offensive Speech Identification and Visualization
Christoph Tillmann, Aashka Trivedi, Sara Rosenthal +4
Offensive language such as hate, abuse, and profanity (HAP) occurs in various content on the web. While previous work has mostly dealt with sentence level annotations, there have b…