From the 1 of 9 linked papers with an AI index.
9 papers
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
Winston Zeng, Ali Emami, Jinho D. Choi
The paper introduces a large inventory of persona vectors to systematically probe open-weight language models, categorizing traits as naturally expressed, steerable, or resistant,…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
If Only My CGM Could Speak: A Privacy-Preserving Agent for Question Answering over Continuous Glucose Data
Yanjun Cui, Ali Emami, Temiloluwa Prioleau +1
Continuous glucose monitors (CGMs) used in diabetes care collect rich personal health data that could improve day-to-day self-management. However, current patient platforms only of…
DART: Mitigating Harm Drift in Difference-Aware LLMs via Distill-Audit-Repair Training
Ziwen Pan, Zihan Liang, Jad Kabbara +1
Large language models (LLMs) tuned for safety often avoid acknowledging demographic differences, even when such acknowledgment is factually correct (e.g., ancestry-based disease in…
Memory Dial: A Training Framework for Controllable Memorization in Language Models
Xiangbo Zhang, Ali Emami
Memorization in language models is widely studied but remains difficult to isolate and control. Understanding when and what models memorize is essential for explaining their predic…
Trace-of-Thought Prompting: Investigating Prompt-Based Knowledge Distillation Through Question Decomposition
Tyler McDonald, Ali Emami
Knowledge distillation allows smaller neural networks to emulate the performance of larger, teacher models with reduced computational demands. Traditional methods for Large Languag…