TinyLlama: An Open-Source Small Language Model
arXiv:2401.02385
Abstract
We present TinyLlama, a compact 1.1B language model pretrained on around 1 trillion tokens for approximately 3 epochs. Building on the architecture and tokenizer of Llama 2, TinyLlama leverages various advances contributed by the open-source community (e.g., FlashAttention and Lit-GPT), achieving better computational efficiency. Despite its relatively small size, TinyLlama demonstrates remarkable performance in a series of downstream tasks. It significantly outperforms existing open-source language models with comparable sizes. Our model checkpoints and code are publicly available on GitHub at https://github.com/jzhang38/TinyLlama.
Technical Report
Cited by in corpus (5)
- Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey
- OpenECAD: An Efficient Visual Language Model for Editable 3D-CAD Design
- Tucano: Advancing Neural Text Generation for Portuguese
- Camera Control at the Edge with Language Models for Scene Understanding
- Demystifying and Improving Lazy Promotion in Cache Eviction