ChatGPT: Jack of all trades, master of none
arXiv:2302.10724 · doi:10.1016/j.inffus.2023.101861
Abstract
OpenAI has released the Chat Generative Pre-trained Transformer (ChatGPT) and revolutionized the approach in artificial intelligence to human-model interaction. Several publications on ChatGPT evaluation test its effectiveness on well-known natural language processing (NLP) tasks. However, the existing studies are mostly non-automated and tested on a very limited scale. In this work, we examined ChatGPT's capabilities on 25 diverse analytical NLP tasks, most of them subjective even to humans, such as sentiment analysis, emotion recognition, offensiveness, and stance detection. In contrast, the other tasks require more objective reasoning like word sense disambiguation, linguistic acceptability, and question answering. We also evaluated GPT-4 model on five selected subsets of NLP tasks. We automated ChatGPT and GPT-4 prompting process and analyzed more than 49k responses. Our comparison of its results with available State-of-the-Art (SOTA) solutions showed that the average loss in quality of the ChatGPT model was about 25% for zero-shot and few-shot evaluation. For GPT-4 model, a loss for semantic tasks is significantly lower than for ChatGPT. We showed that the more difficult the task (lower SOTA performance), the higher the ChatGPT loss. It especially refers to pragmatic NLP problems like emotion recognition. We also tested the ability to personalize ChatGPT responses for selected subjective tasks via Random Contextual Few-Shot Personalization, and we obtained significantly better user-based predictions. Additional qualitative analysis revealed a ChatGPT bias, most likely due to the rules imposed on human trainers by OpenAI. Our results provide the basis for a fundamental discussion of whether the high quality of recent predictive NLP models can indicate a tool's usefulness to society and how the learning and validation procedures for such systems should be established.
preprint
References in corpus (18)
- Training language models to follow instructions with human feedback
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Capabilities of GPT-4 on Medical Challenge Problems
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
- ChatGPT: The End of Online Exam Integrity?
- How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
- Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models
- Instruction Tuning with GPT-4
- News Summarization and Evaluation in the Era of GPT-3
- Entailment as Few-Shot Learner
- On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
- ST-MoE: Designing Stable and Transferable Sparse Expert Models
- Will Affective Computing Emerge from Foundation Models and General AI? A First Evaluation on ChatGPT
- Spam Detection Using BERT
- Targeted Phishing Campaigns using Large Scale Language Models
- Classification of Natural Language Processing Techniques for Requirements Engineering
- MetaQA: Combining Expert Agents for Multi-Skill Question Answering
Cited by in corpus (32)
- Summary of ChatGPT-Related Research and Perspective Towards the Future of Large Language Models
- Artificial muses: Generative Artificial Intelligence Chatbots Have Risen to Human-Level Creativity
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data
- Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
- Factuality Challenges in the Era of Large Language Models
- Towards Interpretable Mental Health Analysis with Large Language Models
- ChatGPT v Bard v Bing v Claude 2 v Aria v human-expert. How good are AI chatbots at scientific writing?
- Can we trust the evaluation on ChatGPT?
- Can ChatGPT evaluate research quality?
- Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness
- ChatGPT in Veterinary Medicine: A Practical Guidance of Generative Artificial Intelligence in Clinics, Education, and Research
- From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language Models
- CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation
- Correctness Comparison of ChatGPT-4, Gemini, Claude-3, and Copilot for Spatial Tasks
- Mining experimental data from Materials Science literature with Large Language Models: an evaluation study
- MetaSegNet: Metadata-collaborative Vision-Language Representation Learning for Semantic Segmentation of Remote Sensing Images
- Evaluating the quality of published medical research with ChatGPT
- Take It, Leave It, or Fix It: Measuring Productivity and Trust in Human-AI Collaboration
- Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
- Can ChatGPT Read Who You Are?
- LLMs as Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines with LLMs
- "Only ChatGPT gets me": An Empirical Analysis of GPT versus other Large Language Models for Emotion Detection in Text
- Can generative AI figure out figurative language? The influence of idioms on essay scoring by ChatGPT, Gemini, and Deepseek
- FlowGPT: Exploring Domains, Output Modalities, and Goals of Community-Generated AI Chatbots
- Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy
- Fine-tuning Strategies for Domain Specific Question Answering under Low Annotation Budget Constraints
- Towards More Robust NLP System Evaluation: Handling Missing Scores in Benchmarks
- Predicting stock prices with ChatGPT-annotated Reddit sentiment
- AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
- Ethical AI prompt recommendations in large language models using collaborative filtering
- Generative AI and the future of scientometrics: current topics and future questions
- Transductive Learning for Textual Few-Shot Classification in API-based Embedding Models