2 papers
cs.CR2026
Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective
Meifang Chen, Zhe Yang, Huang Nianchen +4
Code secrets are sensitive assets for software developers, and their leakage poses significant cybersecurity risks. While the rapid development of AI code assistants powered by Cod…
cs.CL2026
Data Compressibility Quantifies LLM Memorization
Yizhan Huang, Zhe Yang, Meifang Chen +3
Large Language Models (LLMs) are known to memorize portions of their training data, sometimes even reproduce content verbatim when prompted appropriately. Despite substantial inter…