1.6k citations · 1.6k across the 8 of their papers we have counts for
8 papers
LOTOS: Layer-wise Orthogonalization for Training Robust Ensembles
Ali Ebrahimpour-Boroojeny, Hari Sundaram, Varun Chandrasekaran
Transferability of adversarial examples is a well-known property that endangers all classification models, even those that are only accessible through black-box queries. Prior work…
Generative Monoculture in Large Language Models
Fan Wu, Emily Black, Varun Chandrasekaran
We introduce {\em generative monoculture}, a behavior observed in large language models (LLMs) characterized by a significant narrowing of model output diversity relative to availa…
Bypassing LLM Watermarks with Color-Aware Substitutions
Qilong Wu, Varun Chandrasekaran
Watermarking approaches are proposed to identify if text being circulated is human or large language model (LLM) generated. The state-of-the-art watermarking strategy of Kirchenbau…
Teaching Language Models to Hallucinate Less with Synthetic Tasks
Erik Jones, Hamid Palangi, Clarisse Simões +5
Large language models (LLMs) frequently hallucinate on abstractive summarization tasks such as document-based question-answering, meeting summarization, and clinical report generat…
KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval
Marah I Abdin, Suriya Gunasekar, Varun Chandrasekaran +5
We study the ability of state-of-the art models to answer constraint satisfaction queries for information retrieval (e.g., 'a list of ice cream shops in San Diego'). In the past, s…
Why Train More? Effective and Efficient Membership Inference via Memorization
Jihye Choi, Shruti Tople, Varun Chandrasekaran +1
Membership Inference Attacks (MIAs) aim to identify specific data samples within the private training dataset of machine learning models, leading to serious privacy violations and…