1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.CR2024★ 1 cited
Supporting Human Raters with the Detection of Harmful Content using Large Language Models
Kurt Thomas, Patrick Gage Kelley, David Tao +7
In this paper, we explore the feasibility of leveraging large language models (LLMs) to automate or otherwise assist human raters with identifying harmful content including hate sp…
cs.CL2023
RETSim: Resilient and Efficient Text Similarity
Marina Zhang, Owen Vallis, Aysegul Bumin +2
This paper introduces RETSim (Resilient and Efficient Text Similarity), a lightweight, multilingual deep learning model trained to produce robust metric embeddings for near-duplica…