Gemini: A Family of Highly Capable Multimodal Models
arXiv:2312.11805
Abstract
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging from complex reasoning tasks to on-device memory-constrained use-cases. Evaluation on a broad range of benchmarks shows that our most-capable Gemini Ultra model advances the state of the art in 30 of 32 of these benchmarks - notably being the first model to achieve human-expert performance on the well-studied exam benchmark MMLU, and improving the state of the art in every one of the 20 multimodal benchmarks we examined. We believe that the new capabilities of the Gemini family in cross-modal reasoning and language understanding will enable a wide variety of use cases. We discuss our approach toward post-training and deploying Gemini models responsibly to users through services including Gemini, Gemini Advanced, Google AI Studio, and Cloud Vertex AI.
Cited by in corpus (63)
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Unleashing the potential of prompt engineering for large language models
- OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
- Process Modeling With Large Language Models
- Materials science in the era of large language models: a perspective
- A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
- Tool Learning with Large Language Models: A Survey
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
- Learning from models beyond fine-tuning
- Exploring the psychology of LLMs' Moral and Legal Reasoning
- Neural Natural Language Processing for Long Texts: A Survey on Classification and Summarization
- Language Models as Zero-Shot Trajectory Generators
- A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
- Reinforcement Learning for Generative AI: State of the Art, Opportunities and Open Research Challenges
- Interactive Question Answering Systems: Literature Review
- Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
- Multimodal Foundation Models for Material Property Prediction and Discovery
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- Large language models as oracles for instantiating ontologies with domain-specific knowledge
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- "It's Kind of Context Dependent": Understanding Blind and Low Vision People's Video Accessibility Preferences Across Viewing Scenarios
- The Dawn of AI-Native EDA: Opportunities and Challenges of Large Circuit Models
- ProMoAI: Process Modeling with Generative AI
- Transfer Learning with Foundational Models for Time Series Forecasting using Low-Rank Adaptations
- Speed and Conversational Large Language Models: Not All Is About Tokens per Second
- Forecasting high-impact research topics via machine learning on evolving knowledge graphs
- Transforming Agency. On the mode of existence of Large Language Models
- Agents for self-driving laboratories applied to quantum computing
- Large-scale moral machine experiment on large language models
- Large Language Models for Combinatorial Optimization: A Systematic Review
- A Survey on Personalized Content Synthesis with Diffusion Models
- Structured Generative Models for Scene Understanding
- Using AI Large Language Models for Grading in Education: A Hands-On Test for Physics
- Adaptive Intelligence: leveraging insights from adaptive behavior in animals to build flexible AI systems
- Instructor-Worker Large Language Model System for Policy Recommendation: a Case Study on Air Quality Analysis of the January 2025 Los Angeles Wildfires
- LLM-Assisted Visual Analytics: Opportunities and Challenges
- Generative Artificial Intelligence-Guided User Studies: An Application for Air Taxi Services
- Generative AI in Health Economics and Outcomes Research: A Taxonomy of Key Definitions and Emerging Applications, an ISPOR Working Group Report
- Evaluating Company-specific Biases in Financial Sentiment Analysis using Large Language Models
- Mosaic: Composite Projection Pruning for Resource-efficient LLMs
- MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
- Semantic Communications Services within Generalist Operated Networks
- Bine Trees: Enhancing Collective Operations by Optimizing Communication Locality
- Automated Theorem Provers Help Improve Large Language Model Reasoning
- Visual Hindsight Self-Imitation Learning for Interactive Navigation
- MULTI: Multimodal Understanding Leaderboard with Text and Images
- Camera Control at the Edge with Language Models for Scene Understanding
- FANAL -- Financial Activity News Alerting Language Modeling Framework
- On the attribution of confidence to large language models
- Perovskite-R1: a domain-specialized large language model for intelligent discovery of precursor additives and experimental design
- Visual Analysis of LLM-based Entity Resolution from Scientific Papers
- Textual interpretation of transient image classifications from large language models
- Bringing Multi-Modal Multi-Task Federated Foundation Models to Education Domain: Prospects and Challenges
- Mapping Diffuse Radio Sources Using TUNA: A Transformer-Based Deep Learning Approach
- SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
- QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
- PhenoAssistant: A Conversational Multi-Agent AI System for Automated Plant Phenotyping
- A Language-Guided Multimodal Foundation Model for Zero-Shot and Multi-Task Brain Signal Analysis
- WikiHint: A Human-Annotated Dataset for Hint Ranking and Generation
- A Roadmap for Tamed Interactions with Large Language Models
- ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics