4 papers
Improving Self Consistency in LLMs through Probabilistic Tokenization
Ashutosh Sathe, Divyanshu Aggarwal, Sunayana Sitaram
Prior research has demonstrated noticeable performance gains through the use of probabilistic tokenizations, an approach that involves employing multiple tokenizations of the same…
Efficient Training of Language Models with Compact and Consistent Next Token Distributions
Ashutosh Sathe, Sunita Sarawagi
Maximizing the likelihood of the next token is an established, statistically sound objective for pre-training language models. In this paper we show that we can train better models…
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
Prachi Jain, Ashutosh Sathe, Varun Gumma +2
Pretrained Language Models (PLMs) are widely used in NLP for various tasks. Recent studies have identified various biases that such models exhibit and have proposed methods to corr…
Benchmarking and Improving Text-to-SQL Generation under Ambiguity
Adithya Bhaskar, Tushar Tomar, Ashutosh Sathe +1
Research in Text-to-SQL conversion has been largely benchmarked against datasets where each text query corresponds to one correct SQL. However, natural language queries over real-l…