4 papers
Data Distribution as a Lever for Guiding Optimizers Toward Superior Generalization in LLMs
Tushaar Gangavarapu, Jiping Li, Christopher Vattheuer +2
Can modifying the training data distribution guide optimizers toward solutions with improved generalization when training large language models (LLMs)? In this work, we theoretical…
Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models
Son Quoc Tran, Tushaar Gangavarapu, Nicholas Chernogor +2
We often rely on our intuition to anticipate the direction of a conversation. Endowing automated systems with similar foresight can enable them to assist human-human interactions.…
GPT-2 Through the Lens of Vector Symbolic Architectures
Johannes Knittel, Tushaar Gangavarapu, Hendrik Strobelt +1
Understanding the general priniciples behind transformer models remains a complex endeavor. Experiments with probing and disentangling features using sparse autoencoders (SAE) sugg…
MambaByte: Token-free Selective State Space Model
Junxiong Wang, Tushaar Gangavarapu, Jing Nathan Yan +1
Token-free language models learn directly from raw bytes and remove the inductive bias of subword tokenization. Operating on bytes, however, results in significantly longer sequenc…