Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges
arXiv:2412.03220 · doi:10.1109/ACCESS.2024.3482107
Abstract
Large Language Models (LLMs) represent a class of deep learning models adept at understanding natural language and generating coherent responses to various prompts or queries. These models far exceed the complexity of conventional neural networks, often encompassing dozens of neural network layers and containing billions to trillions of parameters. They are typically trained on vast datasets, utilizing architectures based on transformer blocks. Present-day LLMs are multi-functional, capable of performing a range of tasks from text generation and language translation to question answering, as well as code generation and analysis. An advanced subset of these models, known as Multimodal Large Language Models (MLLMs), extends LLM capabilities to process and interpret multiple data modalities, including images, audio, and video. This enhancement empowers MLLMs with capabilities like video editing, image comprehension, and captioning for visual content. This survey provides a comprehensive overview of the recent advancements in LLMs. We begin by tracing the evolution of LLMs and subsequently delve into the advent and nuances of MLLMs. We analyze emerging state-of-the-art MLLMs, exploring their technical features, strengths, and limitations. Additionally, we present a comparative analysis of these models and discuss their challenges, potential limitations, and prospects for future development.
References in corpus (17)
- Generative Adversarial Networks: An Overview
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- Pre-Training with Whole Word Masking for Chinese BERT
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- Revisiting Pre-Trained Models for Chinese Natural Language Processing
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
- OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization
- Survey on Large Language Model-Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods
- Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era
- An Algorithm-Hardware Co-Optimized Framework for Accelerating N:M Sparse Transformers
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Security and Privacy Challenges of Large Language Models: A Survey
- Efficient Fine-Tuning of BERT Models on the Edge
- Unified Low-Resource Sequence Labeling by Sample-Aware Dynamic Sparse Finetuning
- Increasing Model Capacity for Free: A Simple Strategy for Parameter Efficient Fine-tuning