2 papers
cs.LG2025
Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks
Yixin Cheng, Hongcheng Guo, Yangming Li +1
Text watermarking aims to subtly embed statistical signals into text by controlling the Large Language Model (LLM)'s sampling process, enabling watermark detectors to verify that t…
cs.LG2024
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
Yixin Cheng, Markos Georgopoulos, Volkan Cevher +1
Large Language Models (LLMs) are susceptible to Jailbreaking attacks, which aim to extract harmful information by subtly modifying the attack query. As defense mechanisms evolve, d…