activity
20192022
most citedGoogle COVID-19 Search Trends Symptoms Dataset: Anonymization Process Description (version 1.0)

16 citations · 20 across the 3 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL20221 cited

QUILL: Query Intent with Large Language Models using Retrieval Augmentation and Multi-stage Distillation

Krishna Srinivasan, Karthik Raman, Anupam Samanta +3

Large Language Models (LLMs) have shown impressive results on a variety of text understanding tasks. Search queries though pose a unique challenge, given their short-length and lac…

cs.CL20223 cited

FiD-Light: Efficient and Effective Retrieval-Augmented Text Generation

Sebastian Hofstätter, Jiecao Chen, Karthik Raman +1

Retrieval-augmented generation models offer many benefits over standalone language models: besides a textual answer to a given query they provide provenance items retrieved from an…

cs.CL2020

DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries

Aditi Chaudhary, Karthik Raman, Krishna Srinivasan +1

Pre-trained multilingual language models such as mBERT have shown immense gains for several natural language processing (NLP) tasks, especially in the zero-shot cross-lingual setti…

cs.CL2020

DiPair: Fast and Accurate Distillation for Trillion-Scale Text Matching and Pair Modeling

Jiecao Chen, Liu Yang, Karthik Raman +6

Pre-trained models like BERT (Devlin et al., 2018) have dominated NLP / IR applications such as single sentence classification, text pair classification, and question answering. Ho…

cs.CL2019

Learning Multilingual Word Embeddings Using Image-Text Data

Karan Singhal, Karthik Raman, Balder ten Cate

There has been significant interest recently in learning multilingual word embeddings -- in which semantically similar words across languages have similar embeddings. State-of-the-…