3 papers
cs.CL2025
Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation
Collin Zhang, Fei Huang, Chenhan Yuan +1
Large language models (LLMs) often experience language confusion, which is the unintended mixing of languages during text generation. Current solutions to this problem either neces…
cs.LG2025
Harnessing the Universal Geometry of Embeddings
Rishi Jha, Collin Zhang, Vitaly Shmatikov +1
We introduce the first method for translating text embeddings from one vector space to another without any paired data, encoders, or predefined sets of matches. Our unsupervised ap…
cs.CL2025
Universal Zero-shot Embedding Inversion
Collin Zhang, John X. Morris, Vitaly Shmatikov
Embedding inversion, i.e., reconstructing text given its embedding and black-box access to the embedding encoder, is a fundamental problem in both NLP and security. From the NLP pe…