activity
20232026
most citedHateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.CRShow all

5 papers · 1 filter

cs.CR2026

EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models

Yiting Qu, Ziqing Yang, Chi Cui +3

Hidden chain-of-thought (CoT) traces, especially those from frontier proprietary large reasoning models (LRMs), are valuable model assets. Yet whether these hidden CoTs can be dire…

cs.CR2026

Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks

Junjie Chu, Yiting Qu, Ye Leng +4

Large Language Models (LLMs) are increasingly trained to align with human values, primarily focusing on task level, i.e., refusing to execute directly harmful tasks. However, a sub…

cs.CR2025

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions

Yiting Qu, Ziqing Yang, Yihan Ma +3

Recent advances in text-to-image diffusion models have enabled the creation of a new form of digital art: optical illusions--visual tricks that create different perceptions of real…

cs.CR2025

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities

Yiting Qu, Michael Backes, Yang Zhang

Vision-language models (VLMs) are increasingly applied to identify unsafe or inappropriate images due to their internal ethical standards and powerful reasoning abilities. However,…

cs.CR20251 cited

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

Xinyue Shen, Yixin Wu, Yiting Qu +3

Large Language Models (LLMs) have raised increasing concerns about their misuse in generating hate speech. Among all the efforts to address this issue, hate speech detectors play a…