activity
20242026
most citedWhen Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity

1 citations · 1 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CV2026

The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning

Renmiao Chen, Yida Lu, Shiyao Cui +6

As Multimodal Large Language Models (MLLMs) acquire stronger reasoning capabilities to handle complex, multi-image instructions, this advancement may pose new safety risks. We stud…

cs.CL2025

Speculating LLMs' Chinese Training Data Pollution from Their Tokens

Qingjie Zhang, Di Wang, Haoting Qian +7

Tokens are basic elements in the datasets for LLM training. It is well-known that many tokens representing Chinese phrases in the vocabulary of GPT (4o/4o-mini/o1/o3/4.5/4.1/o4-min…

cs.CL2025

Understanding the Dilemma of Unlearning for Large Language Models

Qingjie Zhang, Haoting Qian, Zhicong Huang +5

Unlearning seeks to remove specific knowledge from large language models (LLMs), but its effectiveness remains contested. On one side, "forgotten" knowledge can often be recovered…

cs.CL20251 cited

When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity

Shiyao Cui, Xijia Feng, Yingkang Wang +6

Emojis are globally used non-verbal cues in digital communication, and extensive research has examined how large language models (LLMs) understand and utilize emojis across context…

cs.CL2025

Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings

Shujian Yang, Shiyao Cui, Chuanrui Hu +5

Detecting toxic content using language models is important but challenging. While large language models (LLMs) have demonstrated strong performance in understanding Chinese, recent…

cs.CL2024

Understanding the Dark Side of LLMs' Intrinsic Self-Correction

Qingjie Zhang, Di Wang, Haoting Qian +7

Intrinsic self-correction was proposed to improve LLMs' responses via feedback prompts solely based on their inherent capability. However, recent works show that LLMs' intrinsic se…