Rethinking Search: Making Domain Experts out of Dilettantes
arXiv:2105.02274 · doi:10.1145/3476415.3476428
Abstract
When experiencing an information need, users want to engage with a domain expert, but often turn to an information retrieval system, such as a search engine, instead. Classical information retrieval systems do not answer information needs directly, but instead provide references to (hopefully authoritative) answers. Successful question answering systems offer a limited corpus created on-demand by human experts, which is neither timely nor scalable. Pre-trained language models, by contrast, are capable of directly generating prose that may be responsive to an information need, but at present they are dilettantes rather than domain experts -- they do not have a true understanding of the world, they are prone to hallucinating, and crucially they are incapable of justifying their utterances by referring to supporting documents in the corpus they were trained over. This paper examines how ideas from classical information retrieval and pre-trained language models can be synthesized and evolved into systems that truly deliver on the promise of domain expert advice.
References in corpus (12)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Deep Relevance Matching Model for Ad-hoc Retrieval
- mT5: A massively multilingual pre-trained text-to-text transformer
- End-to-End Neural Ad-hoc Ranking with Kernel Pooling
- Machine Comprehension Using Match-LSTM and Answer Pointer
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Unifying Vision-and-Language Tasks via Text Generation
- Quasar: Datasets for Question Answering by Search and Reading
- Measuring and Reducing Gendered Correlations in Pre-trained Models
- A Compare-Aggregate Model for Matching Text Sequences
- Distilling Dense Representations for Ranking using Tightly-Coupled Teachers
- Leveraging Semantic and Lexical Matching to Improve the Recall of Document Retrieval Systems: A Hybrid Approach
Cited by in corpus (12)
- Information Retrieval: Recent Advances and Beyond
- GERE: Generative Evidence Retrieval for Fact Verification
- CorpusBrain: Pre-train a Generative Retrieval Model for Knowledge-Intensive Language Tasks
- Continual Learning for Generative Retrieval over Dynamic Corpora
- A Unified Generative Retriever for Knowledge-Intensive Language Tasks via Prompt Learning
- Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
- The Role of Complex NLP in Transformers for Text Ranking?
- Advances in Artificial Intelligence: A Review for the Creative Industries
- Control Search Rankings, Control the World: What is a Good Search Engine?
- Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval
- Constrained Auto-Regressive Decoding Constrains Generative Retrieval
- Sponsored Question Answering