75 citations · 77 across the 18 of their papers we have counts for
16 papers · 1 filter
Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation
Junyu Lu, Kaiyuan Liu, Kaichun Wang +8
Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such fe…
Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos
Junyu Lu, Deyi Ji, Liqun Liu +9
Hateful videos have become prevalent on online platforms, highlighting an urgent need for effective detection. However, existing studies primarily focus on binary classification an…
Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes
Weiming Wang, Junyu Lu, Han Wang +5
Research on harmful meme detection has garnered significant attention, resulting in the development of numerous datasets and methods. However, progress in detecting Chinese harmful…
Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis
Junyu Lu, Deyi Ji, Xuanyi Liu +5
Large language models for subjectivity analysis are typically trained with aggregated labels, which compress variations in human judgment into a single supervision signal. This par…
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
Kelaiti Xiao, Liang Yang, Dongyu Zhang +2
We study idiom-based visual puns--images that align an idiom's literal and figurative meanings--and present an iterative framework that coordinates a large language model (LLM), a…
Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks
Zewen Bai, Liang Yang, Shengdi Yin +2
The proliferation of hate speech has inflicted significant societal harm, with its intensity and directionality closely tied to specific targets and arguments. In recent years, num…