4 papers · 1 filter
Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
Shayan Talaei, Abhinav Chinta, Devvrit Khatri +3
Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering user decisions at scale. Such preferential biases can be intro…
Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
Baturay Saglam, Paul Kassianik, Blaine Nelson +3
Understanding the latent space geometry of large language models (LLMs) is key to interpreting their behavior and improving alignment. Yet it remains unclear to what extent LLMs li…
Learning Task Representations from In-Context Learning
Baturay Saglam, Xinyang Hu, Zhuoran Yang +2
Large language models (LLMs) have demonstrated remarkable proficiency in in-context learning (ICL), where models adapt to new tasks through example-based prompts without requiring…
Cognitive Overload Attack:Prompt Injection for Long Context
Bibek Upadhayay, Vahid Behzadan, Amin Karbasi
Large Language Models (LLMs) have demonstrated remarkable capabilities in performing tasks across various domains without needing explicit retraining. This capability, known as In-…