4 papers
LSHBloom: Memory-efficient, Extreme-scale Document Deduplication
Arham Khan, Robert Underwood, Carlo Siebenschuh +7
Contemporary large language model (LLM) training pipelines require the assembly of internet-scale databases full of text data from a variety of sources (e.g., web, academic, and pu…
HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights
Ozan Gokdemir, Carlo Siebenschuh, Alexander Brace +21
The volume of scientific literature is growing exponentially, leading to underutilized discoveries, duplicated efforts, and limited cross-disciplinary collaboration. Retrieval Augm…
Deep Model Merging: The Sister of Neural Network Interpretability -- A Survey
Arham Khan, Todd Nief, Nathaniel Hudson +6
We survey the model merging literature through the lens of loss landscape geometry to connect observations from empirical studies on model merging and loss landscape analysis to ph…
Mitigating Memorization In Language Models
Mansi Sakarvadia, Aswathy Ajith, Arham Khan +6
Language models (LMs) can "memorize" information, i.e., encode training data in their weights in such a way that inference-time queries can lead to verbatim regurgitation of that d…