2 papers
cs.CR2025
Watermarks for Language Models via Probabilistic Automata
Yangkun Wang, Jingbo Shang
A recent watermarking scheme for language models achieves distortion-free embedding and robustness to edit-distance attacks. However, it suffers from limited generation diversity a…
cs.CL2023
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Zi Lin, Zihan Wang, Yongqi Tong +4
Despite remarkable advances that large language models have achieved in chatbots, maintaining a non-toxic user-AI interactive environment has become increasingly critical nowadays.…