4 papers · 1 filter
Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation
Collin Zhang, Fei Huang, Chenhan Yuan +1
Large language models (LLMs) often experience language confusion, which is the unintended mixing of languages during text generation. Current solutions to this problem either neces…
Universal Zero-shot Embedding Inversion
Collin Zhang, John X. Morris, Vitaly Shmatikov
Embedding inversion, i.e., reconstructing text given its embedding and black-box access to the embedding encoder, is a fundamental problem in both NLP and security. From the NLP pe…
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
Collin Zhang, Tingwei Zhang, Vitaly Shmatikov
We design, implement, and evaluate adversarial decoding, a new, generic text generation technique that produces readable documents for different adversarial objectives. Prior metho…
Extracting Prompts by Inverting LLM Outputs
Collin Zhang, John X. Morris, Vitaly Shmatikov
We consider the problem of language model inversion: given outputs of a language model, we seek to extract the prompt that generated these outputs. We develop a new black-box metho…