From the 1 of 9 linked papers with an AI index.
1 citations · 1 across the 5 of their papers we have counts for
7 papers · 1 filter
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
Winston Zeng, Ali Emami, Jinho D. Choi
The paper introduces a large inventory of persona vectors to systematically probe open-weight language models, categorizing traits as naturally expressed, steerable, or resistant,…
DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training
Ziwen Pan, Zihan Liang, Jad Kabbara +1
Large language models (LLMs) tuned for safety often avoid acknowledging demographic differences, even when such acknowledgment is factually correct (e.g., ancestry-based disease in…
Memory Dial: A Training Framework for Controllable Memorization in Language Models
Xiangbo Zhang, Ali Emami
Memorization in language models is widely studied but remains difficult to isolate and control. Understanding when and what models memorize is essential for explaining their predic…
Trace-of-Thought Prompting: Investigating Prompt-Based Knowledge Distillation Through Question Decomposition
Tyler McDonald, Ali Emami
Knowledge distillation allows smaller neural networks to emulate the performance of larger, teacher models with reduced computational demands. Traditional methods for Large Languag…
NYT-Connections: A Deceptively Simple Text Classification Task that Stumps System-1 Thinkers
Angel Yahir Loredo Lopez, Tyler McDonald, Ali Emami
Large Language Models (LLMs) have shown impressive performance on various benchmarks, yet their ability to engage in deliberate reasoning remains questionable. We present NYT-Conne…
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
Robert Morabito, Sangmitra Madhusudan, Tyler McDonald +1
Mitigating explicit and implicit biases in Large Language Models (LLMs) has become a critical focus in the field of natural language processing. However, many current methodologies…