Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Arctic-Embed 2.0: Multilingual Retrieval Without Compromise
Puxuan Yu, Luke Merrick, Gaurav Nuti +1
This paper presents the training methodology of Arctic-Embed 2.0, a set of open-source text embedding models built for accurate and efficient multilingual retrieval. While prior wo…
cs.CL2024
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
Luke Merrick, Danmei Xu, Gaurav Nuti +1
This report describes the training dataset creation and recipe behind the family of \texttt{arctic-embed} text embedding models (a set of five models ranging from 22 to 334 million…