papers

Publications (5)

cs.CL2018

Speech recognition for medical conversations

Chung-Cheng Chiu, Anshuman Tripathi, Katherine Chou +11

In this work we explored building automatic speech recognition models for transcribing doctor patient conversation. We collected a large scale dataset of clinical conversations ($1…

cs.CL2024

Transformer models: an introduction and catalog

Xavier Amatriain, Ananth Sankar, Jie Bing +3

In the past few years we have seen the meteoric appearance of dozens of foundation models of the Transformer family, all of which have memorable and sometimes funny, but not self-e…

cs.CV2020

Smoothed Gaussian Mixture Models for Video Classification and Recommendation

Sirjan Kafle, Aman Gupta, Xue Xia +4

Cluster-and-aggregate techniques such as Vector of Locally Aggregated Descriptors (VLAD), and their end-to-end discriminatively trained equivalents like NetVLAD have recently been…

cs.CL2024

IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages

Mohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar +9

Despite the considerable advancements in English LLMs, the progress in building comparable models for other languages has been hindered due to the scarcity of tailored resources. O…

cs.IR2020

DeText: A Deep Text Ranking Framework with BERT

Weiwei Guo, Xiaowei Liu, Sida Wang +8

Ranking is the most important component in a search system. Mostsearch systems deal with large amounts of natural language data,hence an effective ranking system requires a deep un…