Gemma 2: Improving Open Language Models at a Practical Size
arXiv:2408.00118
Abstract
In this work, we introduce Gemma 2, a new addition to the Gemma family of lightweight, state-of-the-art open models, ranging in scale from 2 billion to 27 billion parameters. In this new version, we apply several known technical modifications to the Transformer architecture, such as interleaving local-global attentions (Beltagy et al., 2020a) and group-query attention (Ainslie et al., 2023). We also train the 2B and 9B models with knowledge distillation (Hinton et al., 2015) instead of next token prediction. The resulting models deliver the best performance for their size, and even offer competitive alternatives to models that are 2-3 times bigger. We release all our models to the community.
Cited by in corpus (11)
- CodEv: An Automated Grading Framework Leveraging Large Language Models for Consistent and Constructive Feedback
- Bridging LMS and generative AI: dynamic course content integration (DCCI) for enhancing student satisfaction and engagement via the ask ME assistant
- Large Language Models for Combinatorial Optimization: A Systematic Review
- Large-scale moral machine experiment on large language models
- Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- Don't Get Too Excited -- Eliciting Emotions in LLMs
- Survey and Evaluation of Converging Architecture in LLMs based on Footsteps of Operations
- Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores
- A Roadmap for Tamed Interactions with Large Language Models
- Can Small Language Models Handle Context-Summarized Multi-Turn Customer-Service QA? A Synthetic Data-Driven Comparative Evaluation