Complex QA and language models hybrid architectures, Survey
arXiv:2302.09051
Abstract
This paper reviews the state-of-the-art of large language models (LLM) architectures and strategies for "complex" question-answering with a focus on hybrid architectures. LLM based chatbot services have allowed anyone to grasp the potential of LLM to solve many common problems, but soon discovered their limitations for complex questions. Addressing more specific, complex questions (e.g., "What is the best mix of power-generation methods to reduce climate change ?") often requires specialized architectures, domain knowledge, new skills, decomposition and multi-step resolution, deep reasoning, sensitive data protection, explainability, and human-in-the-loop processes. Therefore, we review: (1) necessary skills and tasks for handling complex questions and common LLM limits to overcome; (2) dataset, cost functions and evaluation metrics for measuring and improving (e.g. accuracy, explainability, fairness, robustness, groundedness, faithfulness, toxicity...); (3) family of solutions to overcome LLM limitations by (a) training and reinforcement (b) hybridization, (c) prompting, (d) agentic-architectures (agents, tools) and extended reasoning.
References in corpus (126)
- A Simple Framework for Contrastive Learning of Visual Representations
- Survey of Hallucination in Natural Language Generation
- Training language models to follow instructions with human feedback
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Bootstrap your own latent: A new approach to self-supervised Learning
- LoRA: Low-Rank Adaptation of Large Language Models
- BERTScore: Evaluating Text Generation with BERT
- PaLM: Scaling Language Modeling with Pathways
- Improved Baselines with Momentum Contrastive Learning
- Cross-lingual Language Model Pretraining
- Evaluating Large Language Models Trained on Code
- Large Language Models are Zero-Shot Reasoners
- Emergent Abilities of Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Training Compute-Optimal Large Language Models
- A Survey on Active Learning and Human-in-the-Loop Deep Learning for Medical Image Analysis
- ReAct: Synergizing Reasoning and Acting in Language Models
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- The Goldilocks Principle: Reading Children's Books with Explicit Memory Representations
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- BARTScore: Evaluating Generated Text as Text Generation
- Improving language models by retrieving from trillions of tokens
- GLM-130B: An Open Bilingual Pre-trained Model
- Solving Quantitative Reasoning Problems with Language Models
- Galactica: A Large Language Model for Science
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
- ChatGPT is not all you need. A State of the Art Review of large Generative AI models
- Automatic Chain of Thought Prompting in Large Language Models
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- Atlas: Few-shot Learning with Retrieval Augmented Language Models
- A Neural Network Solves, Explains, and Generates University Math Problems by Program Synthesis and Few-Shot Learning at Human Level
- Augmented Language Models: a Survey
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Scalable agent alignment via reward modeling: a research direction
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- STaR: Bootstrapping Reasoning With Reasoning
- The Cost of Training NLP Models: A Concise Overview
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- Prompting Is Programming: A Query Language for Large Language Models
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning
- PAL: Program-aided Language Models
- Decomposed Prompting: A Modular Approach for Solving Complex Tasks
- UL2: Unifying Language Learning Paradigms
- OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization
- Privacy-Preserving Machine Learning: Methods, Challenges and Directions
- Complexity-Based Prompting for Multi-Step Reasoning
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
- MCScript: A Novel Dataset for Assessing Machine Comprehension Using Script Knowledge
- Selective Annotation Makes Language Models Better Few-Shot Learners
- The Many Dimensions of Truthfulness: Crowdsourcing Misinformation Assessments on a Multidimensional Scale
- Language Models are Multilingual Chain-of-Thought Reasoners
- Teaching language models to support answers with verified quotes
- Intermediate-Task Transfer Learning with Pretrained Models for Natural Language Understanding: When and Why Does It Work?
- Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP
- The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning
- Generated Knowledge Prompting for Commonsense Reasoning
- Promptagator: Few-shot Dense Retrieval From 8 Examples
- Active Prompting with Chain-of-Thought for Large Language Models
- Natural Language Understanding with Privacy-Preserving BERT
- PEER: A Collaborative Language Model
- Data Governance in the Age of Large-Scale Data-Driven Language Technology
- A Survey on Automated Fact-Checking
- Beyond Goldfish Memory: Long-Term Open-Domain Conversation
- Confident Adaptive Language Modeling
- Compositional Semantic Parsing with Large Language Models
- Faithful Reasoning Using Large Language Models
- FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts
- GODEL: Large-Scale Pre-Training for Goal-Directed Dialog
- LOREN: Logic-Regularized Reasoning for Interpretable Fact Verification
- FedMatch: Federated Learning Over Heterogeneous Question Answering Data
- GPL: Generative Pseudo Labeling for Unsupervised Domain Adaptation of Dense Retrieval
- Language Models are General-Purpose Interfaces
- InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval
- Chain of Hindsight Aligns Language Models with Feedback
- Recurrent Memory Transformer
- Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models
- Teaching Algorithmic Reasoning via In-context Learning
- Mind's Eye: Grounded Language Model Reasoning through Simulation
- Language Model Cascades
- Offsite-Tuning: Transfer Learning without Full Model
- PINTO: Faithful Language Reasoning Using Prompt-Generated Rationales
- Language Models that Seek for Knowledge: Modular Search & Generation for Dialogue and Prompt Completion
- Teacher-Student Architecture for Knowledge Learning: A Survey
- Multi-hop Question Answering
- Language Models Can Teach Themselves to Program Better
- Answer-Me: Multi-Task Open-Vocabulary Visual Question Answering
- Automatic Evaluation and Moderation of Open-domain Dialogue Systems
- HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System
- Towards Explainable Evaluation Metrics for Natural Language Generation
- Structured Prompting: Scaling In-Context Learning to 1,000 Examples
- COCO-DR: Combating Distribution Shifts in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust Learning
- Mastering the Game of No-Press Diplomacy via Human-Regularized Reinforcement Learning and Planning
- Towards Human Centered AutoML
- Unified Question Generation with Continual Lifelong Learning
- Multimodal Analogical Reasoning over Knowledge Graphs
- Can Open-Domain QA Reader Utilize External Knowledge Efficiently like Humans?
- Knowledge Inheritance for Pre-trained Language Models
- Memory Augmented Large Language Models are Computationally Universal
- HyperTuning: Toward Adapting Large Language Models without Back-propagation
- Task Ambiguity in Humans and Language Models
- CLEVRER-Humans: Describing Physical and Causal Events the Human Way
- PERFECT: Prompt-free and Efficient Few-shot Learning with Language Models
- Composing Ensembles of Pre-trained Models via Iterative Consensus
- Exploring Universal Intrinsic Task Subspace via Prompt Tuning
- The Expertise Problem: Learning from Specialized Feedback
- Retrieval-Augmented Reinforcement Learning
- VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges
- reStructured Pre-training
- Chain of Thought Imitation with Procedure Cloning
- ArchivalQA: A Large-scale Benchmark Dataset for Open Domain Question Answering over Historical News Collections
- Learning Diverse Document Representations with Deep Query Interactions for Dense Retrieval
- MAQA: A Multimodal QA Benchmark for Negation
- A Survey on Table-and-Text HybridQA: Concepts, Methods, Challenges and Future Directions
- Smooth Sailing: Improving Active Learning for Pre-trained Language Models with Representation Smoothness Analysis
- Pre-computed memory or on-the-fly encoding? A hybrid approach to retrieval augmentation makes the most of your compute
- AcTune: Uncertainty-aware Active Self-Training for Semi-Supervised Active Learning with Pretrained Language Models
- Autoencoding Language Model Based Ensemble Learning for Commonsense Validation and Explanation
- Asking for Knowledge: Training RL Agents to Query External Knowledge Using Language
- Teacher Guided Training: An Efficient Framework for Knowledge Transfer
- Spending Thinking Time Wisely: Accelerating MCTS with Virtual Expansions
- Reduce, Reuse, Recycle: Improving Training Efficiency with Distillation