4 papers
Arcee Trinity Large Technical Report
Varun Singh, Lucas Krauss, Sami Jaghouar +23
We present the technical report for Arcee Trinity Large, a sparse Mixture-of-Experts model with 400B total parameters and 13B activated per token. Additionally, we report on Trinit…
INTELLECT-1 Technical Report
Sami Jaghouar, Jack Min Ong, Manveer Basra +9
In this report, we introduce INTELLECT-1, the first 10 billion parameter language model collaboratively trained across the globe, demonstrating that large-scale model training is n…
Domain Adaptation of Llama3-70B-Instruct through Continual Pre-Training and Model Merging: A Comprehensive Evaluation
Shamane Siriwardhana, Mark McQuade, Thomas Gauthier +8
We conducted extensive experiments on domain adaptation of the Meta-Llama-3-70B-Instruct model on SEC data, exploring its performance on both general and domain-specific benchmarks…
Spectrum: Targeted Training on Signal to Noise Ratio
Eric Hartford, Lucas Atkins, Fernando Fernandes Neto +1
Efficiently post-training large language models remains a challenging task due to the vast computational resources required. We present Spectrum, a method that accelerates LLM trai…