The Falcon Series of Open Language Models
arXiv:2311.16867
Abstract
We introduce the Falcon series: 7B, 40B, and 180B parameters causal decoder-only models trained on a diverse high-quality corpora predominantly assembled from web data. The largest model, Falcon-180B, has been trained on over 3.5 trillion tokens of text--the largest openly documented pretraining run. Falcon-180B significantly outperforms models such as PaLM or Chinchilla, and improves upon concurrently developed models such as LLaMA 2 or Inflection-1. It nears the performance of PaLM-2-Large at a reduced pretraining and inference cost, making it, to our knowledge, one of the three best language models in the world along with GPT-4 and PaLM-2-Large. We report detailed evaluations, as well as a deep dive into the methods and custom tooling employed to pretrain Falcon. Notably, we report on our custom distributed training codebase, allowing us to efficiently pretrain these models on up to 4,096 A100s on cloud AWS infrastructure with limited interconnect. We release a 600B tokens extract of our web dataset, as well as the Falcon-7/40/180B models under a permissive license to foster open-science and accelerate the development of an open ecosystem of large language models.
Cited by in corpus (15)
- Large Language Models (LLMs) as Agents for Augmented Democracy
- Text Clustering with Large Language Model Embeddings
- Mask-guided BERT for Few Shot Text Classification
- Leveraging Large Language Models for Patient Engagement: The Power of Conversational AI in Digital Health
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- Isolating Compiler Bugs by Generating Effective Witness Programs with Large Language Models
- Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
- Discrete Prompt Compression with Reinforcement Learning
- CuentosIE: can a chatbot about "tales with a message" help to teach emotional intelligence?
- ColBERT Retrieval and Ensemble Response Scoring for Language Model Question Answering
- Evaluating the Reliability of Self-Explanations in Large Language Models
- InfoTech Assistant: A Multimodal Conversational Agent for InfoTechnology Web Portal Queries
- Don't Get Too Excited -- Eliciting Emotions in LLMs
- Tailored-LLaMA: Optimizing Few-Shot Learning in Pruned LLaMA Models with Task-Specific Prompts
- Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory