2 papers
cs.IR2025
EnronQA: Towards Personalized RAG over Private Documents
Michael J. Ryan, Danmei Xu, Chris Nivera +1
Retrieval Augmented Generation (RAG) has become one of the most popular methods for bringing knowledge-intensive context to large language models (LLM) because of its ability to br…
cs.CL2024
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
Luke Merrick, Danmei Xu, Gaurav Nuti +1
This report describes the training dataset creation and recipe behind the family of \texttt{arctic-embed} text embedding models (a set of five models ranging from 22 to 334 million…