8 papers
AGC-Bench: Measuring Artificial General Creativity
Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai +9
Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from general intelligence. Both ques…
A Pre-Training Analogue of Grokking in Language Models: Tracing Delayed Grammatical Generalization
Sherin Muckatira, Namrata Shivagunde, Vijeta Deshpande +1
Grokking, the phenomenon in which neural networks generalize long after fitting their training data, has been studied in supervised settings on many epochs. LLM pre-training instea…
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
Vijeta Deshpande, Tootiya Giyahchi, Veena Padmanabhan +2
Safety detection models require examples of HHH (Helpful, Harmless, Honest)-violating outputs for robust generalization, however such examples are scarce. Activation Steering (AS)…
Playing with Words, Improving with Rewards: Training Language Models for Creative Association
Vijeta Deshpande, Namrata Shivagunde, Sherin Muckatira +5
Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLMs need to be creative. Yet t…
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
Namrata Shivagunde, Vijeta Deshpande, Sherin Muckatira +1
Pre-training large language models is dominated by the memory cost of storing full-rank weights, gradients, and optimizer states. Low-rank pre-training has emerged to address this,…
Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models
Vijeta Deshpande, Debasmita Ghose, John D. Patterson +2
Diverse language model responses are crucial for creative generation, open-ended tasks, and self-improvement training. We show that common diversity metrics, and even reward models…