3 papers
cs.CL2026
New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
Shiyao Cui, QingLin Zhang, Di Wang +7
Neologisms, emerging terms in meaning or form, can serve as new vehicles for toxic expression, like "country girl" as a stigmatizing label targeting feminism. Such toxic neologisms…
cs.CL2025
Speculating LLMs' Chinese Training Data Pollution from Their Tokens
Qingjie Zhang, Di Wang, Haoting Qian +7
Tokens are basic elements in the datasets for LLM training. It is well-known that many tokens representing Chinese phrases in the vocabulary of GPT (4o/4o-mini/o1/o3/4.5/4.1/o4-min…
cs.CL2025
Understanding the Dark Side of LLMs' Intrinsic Self-Correction
Qingjie Zhang, Di Wang, Haoting Qian +7
Intrinsic self-correction was proposed to improve LLMs' responses via feedback prompts solely based on their inherent capability. However, recent works show that LLMs' intrinsic se…