A Survey of Large Language Models
arXiv:2303.18223 · doi:10.1007/s11704-026-60308-3
Abstract
Language is essentially a complex, intricate system of human expressions governed by grammatical rules. It poses a significant challenge to develop capable AI algorithms for comprehending and grasping a language. As a major approach, language modeling has been widely studied for language understanding and generation in the past two decades, evolving from statistical language models to neural language models. Recently, pre-trained language models (PLMs) have been proposed by pre-training Transformer models over large-scale corpora, showing strong capabilities in solving various NLP tasks. Since researchers have found that model scaling can lead to performance improvement, they further study the scaling effect by increasing the model size to an even larger size. Interestingly, when the parameter scale exceeds a certain level, these enlarged language models not only achieve a significant performance improvement but also show some special abilities that are not present in small-scale language models. To discriminate the difference in parameter scale, the research community has coined the term large language models (LLM) for the PLMs of significant size. Recently, the research on LLMs has been largely advanced by both academia and industry, and a remarkable progress is the launch of ChatGPT, which has attracted widespread attention from society. The technical evolution of LLMs has been making an important impact on the entire AI community, which would revolutionize the way how we develop and use AI algorithms. In this survey, we review the recent advances of LLMs by introducing the background, key findings, and mainstream techniques. In particular, we focus on four major aspects of LLMs, namely pre-training, adaptation tuning, utilization, and capacity evaluation. Besides, we also summarize the available resources for developing LLMs and discuss the remaining issues for future directions.
ongoing work; 144 pages, 1081 citations
Cited by in corpus (99)
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly
- Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
- A Survey on Multimodal Large Language Models
- TALLRec: An Effective and Efficient Tuning Framework to Align Large Language Model with Recommendation
- Recommender Systems in the Era of Large Language Models (LLMs)
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- Large language models in medicine: the potentials and pitfalls
- Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
- Uncovering ChatGPT's Capabilities in Recommender Systems
- A Survey of Knowledge Tracing: Models, Variants, and Applications
- Is ChatGPT Fair for Recommendation? Evaluating Fairness in Large Language Model Recommendation
- Prompt engineering paradigms for medical applications: scoping review and recommendations for better practices
- Large Language Models as Zero-Shot Conversational Recommenders
- PubMed and Beyond: Biomedical Literature Search in the Age of Artificial Intelligence
- Generative AI for Synthetic Data Across Multiple Medical Modalities: A Systematic Review of Recent Developments and Challenges
- How understanding large language models can inform the use of ChatGPT in physics education
- A Scoping Review of ChatGPT Research in Accounting and Finance
- A Study on the Implementation of Generative AI Services Using an Enterprise Data-Based LLM Application Architecture
- Tool Learning with Large Language Models: A Survey
- Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language Models
- UAVs Meet LLMs: Overviews and Perspectives Toward Agentic Low-Altitude Mobility
- The Dawn of Natural Language to SQL: Are We Fully Ready?
- Learning from models beyond fine-tuning
- A Comprehensive Survey on Integrating Large Language Models with Knowledge-Based Methods
- Emotion Detection for Misinformation: A Review
- ChatGPT Informed Graph Neural Network for Stock Movement Prediction
- Language Models as Zero-Shot Trajectory Generators
- Knowledge graph enhanced retrieval-augmented generation for failure mode and effects analysis
- How Good is ChatGPT at Face Biometrics? A First Look into Recognition, Soft Biometrics, and Explainability
- Large Process Models: A Vision for Business Process Management in the Age of Generative AI
- HPC-Coder: Modeling Parallel Programs using Large Language Models
- Fine-tuning and Utilization Methods of Domain-specific LLMs
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- Exploring the Roles of Large Language Models in Reshaping Transportation Systems: A Survey, Framework, and Roadmap
- Achieving Peak Performance for Large Language Models: A Systematic Review
- A Unified Industrial Large Knowledge Model Framework in Industry 4.0 and Smart Manufacturing
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- GeoCode-GPT: A Large Language Model for Geospatial Code Generation Tasks
- Automated scholarly paper review: Concepts, technologies, and challenges
- Several categories of Large Language Models (LLMs): A Short Survey
- Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support
- Large language models for automated scholarly paper review: A survey
- Computational Politeness in Natural Language Processing: A Survey
- Large Knowledge Model: Perspectives and Challenges
- An Unified Search and Recommendation Foundation Model for Cold-Start Scenario
- T-FREX: A Transformer-based Feature Extraction Method from Mobile App Reviews
- Safety Analysis in the Era of Large Language Models: A Case Study of STPA using ChatGPT
- Exploring Gen-AI applications in building research and industry: A review
- Transition Role of Entangled Data in Quantum Machine Learning
- Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
- Potential Benefits of Employing Large Language Models in Research in Moral Education and Development
- Utilizing Cognitive Signals Generated during Human Reading to Enhance Keyphrase Extraction from Microblogs
- ESM-NBR: fast and accurate nucleic acid-binding residue prediction via protein language model feature representation and multi-task learning
- SoVAR: Building Generalizable Scenarios from Accident Reports for Autonomous Driving Testing
- Geo-FuB: A Method for Constructing an Operator-Function Knowledge Base for Geospatial Code Generation Tasks Using Large Language Models
- Computational Argumentation-based Chatbots: a Survey
- Foundation Models and Transformers for Anomaly Detection: A Survey
- Large Language Models at Work in China's Labor Market
- Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field
- MMSR: Symbolic Regression is a Multi-Modal Information Fusion Task
- Supersonic: Learning to Generate Source Code Optimizations in C/C++
- Engineering Artificial Intelligence: Framework, Challenges, and Future Direction
- Automating Traffic Model Enhancement with AI Research Agent
- Low-resource finetuning of foundation models beats state-of-the-art in histopathology
- Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?
- Zero-shot Bilingual App Reviews Mining with Large Language Models
- A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness
- Metamorphic Malware Evolution: The Potential and Peril of Large Language Models
- On Sarcasm Detection with OpenAI GPT-based Models
- Entity Recognition from Colloquial Text
- Emergency Department Decision Support using Clinical Pseudo-notes
- Ultra-Efficient On-Device Object Detection on AI-Integrated Smart Glasses with TinyissimoYOLO
- DeepPatent2: A Large-Scale Benchmarking Corpus for Technical Drawing Understanding
- Generating Universal Adversarial Perturbations for Quantum Classifiers
- Ensemble Neural Networks for Remaining Useful Life (RUL) Prediction
- GPT-DETOX: An In-Context Learning-Based Paraphraser for Text Detoxification
- Generative AI in Health Economics and Outcomes Research: A Taxonomy of Key Definitions and Emerging Applications, an ISPOR Working Group Report
- TDML -- A Trustworthy Distributed Machine Learning Framework
- Foundation Models for the Digital Twin Creation of Cyber-Physical Systems
- Visual Hindsight Self-Imitation Learning for Interactive Navigation
- LLM-ProS: Analyzing Large Language Models' Performance in Competitive Problem Solving
- Exploring the Effectiveness of Abstract Syntax Tree Patterns for Algorithm Recognition
- TickIt: Leveraging Large Language Models for Automated Ticket Escalation
- Comprehensive Modeling and Question Answering of Cancer Clinical Practice Guidelines using LLMs
- GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
- ROSGPT_Vision: Commanding Robots Using Only Language Models' Prompts
- Classification of integers based on residue classes via modern deep learning algorithms
- Parameter Efficient Diverse Paraphrase Generation Using Sequence-Level Knowledge Distillation
- Hybrid Quantum-inspired Resnet and Densenet for Pattern Recognition
- Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques
- CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
- On the Generalization and Adaptation Ability of Machine-Generated Text Detectors in Academic Writing
- Developing Artificial Mechanics Intuitions from Extremely Small Data
- Generative Agents Navigating Digital Libraries
- A Computational Approach to Modeling Conversational Systems: Analyzing Large-Scale Quasi-Patterned Dialogue Flows
- Continually Learn to Map Visual Concepts to Large Language Models in Resource-constrained Environments
- Data2Concept2Text: An Explainable Multilingual Framework for Data Analysis Narration
- Large Language Models -- the Future of Fundamental Physics?