5 papers
Fara-1.5: Scalable Learning Environments for Computer Use Agents
Ahmed Awadallah, Sahil Gupta, Yash Lara +12
Collecting computer use data from human demonstrations is expensive and slow, motivating the need for scalable generation strategies. This requires two key ingredients: environment…
NOMAD Projection
Brandon Duderstadt, Zach Nussbaum, Laurens van der Maaten
The rapid adoption of generative AI has driven an explosion in the size of datasets consumed and produced by AI models. Traditional methods for unstructured data visualization, suc…
Training Sparse Mixture Of Experts Text Embedding Models
Zach Nussbaum, Brandon Duderstadt
Transformer-based text embedding models have improved their performance on benchmarks like MIRACL and BEIR by increasing their parameter counts. However, this scaling approach intr…
CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking
Tarun Suresh, Revanth Gangi Reddy, Yifei Xu +4
Effective code retrieval plays a crucial role in advancing code generation, bug fixing, and software maintenance, particularly as software systems increase in complexity. While cur…
Nomic Embed: Training a Reproducible Long Context Text Embedder
Zach Nussbaum, John X. Morris, Brandon Duderstadt +1
This technical report describes the training of nomic-embed-text-v1, the first fully reproducible, open-source, open-weights, open-data, 8192 context length English text embedding…