1 citations · 1 across the 8 of their papers we have counts for
5 papers · 1 filter
Continual Learning via Sparse Memory Finetuning
Jessy Lin, Luke Zettlemoyer, Gargi Ghosh +4
Modern language models are powerful, but typically static after deployment. A major obstacle to building models that continually learn over time is catastrophic forgetting, where u…
Detecting Prefix Bias in LLM-based Reward Models
Ashwin Kumar, Yuzi He, Aram H. Markosyan +2
Reinforcement Learning with Human Feedback (RLHF) has emerged as a key paradigm for task-specific fine-tuning of language models using human preference data. While numerous publicl…
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu +9
As large language models (LLMs) become increasingly prevalent across many real-world applications, understanding and enhancing their robustness to adversarial attacks is of paramou…
Using Captum to Explain Generative Language Models
Vivek Miglani, Aobo Yang, Aram H. Markosyan +2
Captum is a comprehensive library for model explainability in PyTorch, offering a range of methods from the interpretability literature to enhance users' understanding of PyTorch m…
Tell Your Story: Task-Oriented Dialogs for Interactive Content Creation
Satwik Kottur, Seungwhan Moon, Aram H. Markosyan +3
People capture photos and videos to relive and share memories of personal significance. Recently, media montages (stories) have become a popular mode of sharing these memories due…