Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
arXiv:2403.05530
Abstract
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio. The family includes two new models: (1) an updated Gemini 1.5 Pro, which exceeds the February version on the great majority of capabilities and benchmarks; (2) Gemini 1.5 Flash, a more lightweight variant designed for efficiency with minimal regression in quality. Gemini 1.5 models achieve near-perfect recall on long-context retrieval tasks across modalities, improve the state-of-the-art in long-document QA, long-video QA and long-context ASR, and match or surpass Gemini 1.0 Ultra's state-of-the-art performance across a broad set of benchmarks. Studying the limits of Gemini 1.5's long-context ability, we find continued improvement in next-token prediction and near-perfect retrieval (>99%) up to at least 10M tokens, a generational leap over existing models such as Claude 3.0 (200k) and GPT-4 Turbo (128k). Finally, we highlight real-world use cases, such as Gemini 1.5 collaborating with professionals on completing their tasks achieving 26 to 75% time savings across 10 different job categories, as well as surprising new capabilities of large language models at the frontier; when given a grammar manual for Kalamang, a language with fewer than 200 speakers worldwide, the model learns to translate English to Kalamang at a similar level to a person who learned from the same content.
Cited by in corpus (26)
- Materials science in the era of large language models: a perspective
- LLM for SoC Security: A Paradigm Shift
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- Playing with words: Comparing the vocabulary and lexical diversity of ChatGPT and humans
- Large language models for automated scholarly paper review: A survey
- Reasoning Beyond Limits: Advances and Open Problems for LLMs
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- Computational Argumentation-based Chatbots: a Survey
- CodEv: An Automated Grading Framework Leveraging Large Language Models for Consistent and Constructive Feedback
- Securing RAG: A Risk Assessment and Mitigation Framework
- Bridging LMS and generative AI: dynamic course content integration (DCCI) for enhancing student satisfaction and engagement via the ask ME assistant
- A recent evaluation on the performance of LLMs on radiation oncology physics using questions of randomly shuffled options
- Large Language Models for Combinatorial Optimization: A Systematic Review
- Large-scale moral machine experiment on large language models
- RepairBench: Leaderboard of Frontier Models for Program Repair
- Predicting Affective States from Screen Text Sentiment
- Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization
- MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
- Sentiment analysis of preservice teachers' reflections using a large language model
- EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
- Sample Efficient Reinforcement Learning via Large Vision Language Model Distillation
- Towards Adaptive Context Management for Intelligent Conversational Question Answering
- Wrong Answers Can Also Be Useful: PlausibleQA -- A Large-Scale QA Dataset with Answer Plausibility Scores
- GenCeption: Evaluate Vision LLMs with Unlabeled Unimodal Data
- A Computational Approach to Modeling Conversational Systems: Analyzing Large-Scale Quasi-Patterned Dialogue Flows
- A Roadmap for Tamed Interactions with Large Language Models