31 citations · 61 across the 13 of their papers we have counts for
9 papers · 1 filter
Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation
Junyu Lu, Kaiyuan Liu, Kaichun Wang +8
Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such fe…
Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos
Junyu Lu, Deyi Ji, Liqun Liu +9
Hateful videos have become prevalent on online platforms, highlighting an urgent need for effective detection. However, existing studies primarily focus on binary classification an…
Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes
Weiming Wang, Junyu Lu, Han Wang +5
Research on harmful meme detection has garnered significant attention, resulting in the development of numerous datasets and methods. However, progress in detecting Chinese harmful…
Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting
Jingyi Kang, Junyu Lu, Bo Xu +4
Large language models (LLMs) require robust toxicity evaluation beyond explicit wording. This setting remains underexplored in Chinese, where toxicity may combine semantic indirect…
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
Junyu Lu, Kai Ma, Kaichun Wang +5
Large Language Models (LLMs) have become essential for offensive language detection, yet their ability to handle annotation disagreement remains underexplored. Disagreement samples…
Towards Comprehensive Detection of Chinese Harmful Memes
Junyu Lu, Bo Xu, Xiaokun Zhang +5
This paper has been accepted in the NeurIPS 2024 D & B Track. Harmful memes have proliferated on the Chinese Internet, while research on detecting Chinese harmful memes significant…