2 papers
cs.CL2026
Detoxification for LLM: From Dataset Itself
Wei Shao, Yihang Wang, Gaoyu Zhu +4
Existing detoxification methods for large language models mainly focus on post-training stage or inference time, while few tackle the source of toxicity, namely, the dataset itself…
cs.CL2025
QUITO-X: A New Perspective on Context Compression from the Information Bottleneck Theory
Yihang Wang, Xu Huang, Bowen Tian +6
Generative LLM have achieved remarkable success in various industrial applications, owing to their promising In-Context Learning capabilities. However, the issue of long context in…