activity
20242026
most citedDeMod: A Holistic Tool with Explainable Detection and Personalized Modification for Toxicity Censorship

6 citations · 10 across the 31 of their papers we have counts for

collaborators
Showing 2025Show all

15 papers · 1 filter

cs.LG2025

Metis: Training LLMs with FP4 Quantization

Hengjie Cao, Mengyi Chen, Yifeng Yang +13

This work identifies anisotropy in the singular value spectra of parameters, activations, and gradients as the fundamental barrier to low-bit training of large language models (LLM…

cs.CL2025

"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth

Yaqiong Li, Peng Zhang, Lin Wang +4

Risk perception is subjective, and youth's understanding of toxic content differs from that of adults. Although previous research has conducted extensive studies on toxicity detect…

cs.CL2025

IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization

Yuzhuo Bai, Shitong Duan, Muhua Huang +7

Trained on various human-authored corpora, Large Language Models (LLMs) have demonstrated a certain capability of reflecting specific human-like traits (e.g., personality or values…

cs.AI2025

MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions

Yanxu Zhu, Shitong Duan, Xiangxu Zhang +7

Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despit…

cs.IR2025

Bidirectional Knowledge Distillation for Enhancing Sequential Recommendation with Large Language Models

Jiongran Wu, Jiahao Liu, Dongsheng Li +7

Large language models (LLMs) have demonstrated exceptional performance in understanding and generating semantic patterns, making them promising candidates for sequential recommenda…

cs.IR20251 cited

LLM-Based User Simulation for Low-Knowledge Shilling Attacks on Recommender Systems

Shengkang Gu, Jiahao Liu, Dongsheng Li +7

Recommender systems (RS) are increasingly vulnerable to shilling attacks, where adversaries inject fake user profiles to manipulate system outputs. Traditional attack strategies of…