Gemma: Open Models Based on Gemini Research and Technology
arXiv:2403.08295
Abstract
This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models. Gemma models demonstrate strong performance across academic benchmarks for language understanding, reasoning, and safety. We release two sizes of models (2 billion and 7 billion parameters), and provide both pretrained and fine-tuned checkpoints. Gemma outperforms similarly sized open models on 11 out of 18 text-based tasks, and we present comprehensive evaluations of safety and responsibility aspects of the models, alongside a detailed description of model development. We believe the responsible release of LLMs is critical for improving the safety of frontier models, and for enabling the next wave of LLM innovations.
Cited by in corpus (13)
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- Large language models for automated scholarly paper review: A survey
- OpenECAD: An Efficient Visual Language Model for Editable 3D-CAD Design
- Multi-step retrieval and reasoning improves radiology question answering with large language models
- Large-scale moral machine experiment on large language models
- Helpful assistant or fruitful facilitator? Investigating how personas affect language model behavior
- Helping LLMs Improve Code Generation Using Feedback from Testing and Static Analysis
- Mosaic: Composite Projection Pruning for Resource-efficient LLMs
- Verifying the Robustness of Automatic Credibility Assessment
- AutoMathKG: The automated mathematical knowledge graph based on LLM and vector database
- Evaluating the Reliability of Self-Explanations in Large Language Models
- CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
- A Roadmap for Tamed Interactions with Large Language Models