activity
20232025
most citedBaseline Defenses for Adversarial Attacks Against Aligned Language Models

37 citations · 61 across the 10 of their papers we have counts for

collaborators

10 papers

cs.AI2025

Has My System Prompt Been Used? Large Language Model Prompt Membership Inference

Roman Levin, Valeriia Cherepanova, Abhimanyu Hans +2

Prompt engineering has emerged as a powerful technique for optimizing large language models (LLMs) for specific applications, enabling faster prototyping and improved performance,…

cs.LG20251 cited

Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges

Nayoung Lee, Ziyang Cai, Avi Schwarzschild +2

Large language models often struggle with length generalization and solving complex problem instances beyond their training distribution. We present a self-improvement approach whe…

cs.LG2024

The CLRS-Text Algorithmic Reasoning Language Benchmark

Larisa Markeeva, Sean McLeish, Borja Ibarz +7

Eliciting reasoning capabilities from language models (LMs) is a critical direction on the path towards building intelligent systems. Most recent studies dedicated to reasoning foc…

cs.AI2024

Benchmarking ChatGPT on Algorithmic Reasoning

Sean McLeish, Avi Schwarzschild, Tom Goldstein

We evaluate ChatGPT's ability to solve algorithm problems from the CLRS benchmark suite that is designed for GNNs. The benchmark requires the use of a specified classical algorithm…

cs.LG20245 cited

TOFU: A Task of Fictitious Unlearning for LLMs

Pratyush Maini, Zhili Feng, Avi Schwarzschild +2

Large language models trained on massive corpora of data from the web can memorize and reproduce sensitive or private data raising both legal and ethical concerns. Unlearning, or t…

cs.CL202314 cited

NEFTune: Noisy Embeddings Improve Instruction Finetuning

Neel Jain, Ping-yeh Chiang, Yuxin Wen +10

We show that language model finetuning can be improved, sometimes dramatically, with a simple augmentation. NEFTune adds noise to the embedding vectors during training. Standard fi…