activity
20192023
most citedBenchmarking LLM powered Chatbots: Methods and Metrics

17 citations · 22 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL202317 cited

Benchmarking LLM powered Chatbots: Methods and Metrics

Debarag Banerjee, Pooja Singh, Arjun Avadhanam +1

Autonomous conversational agents, i.e. chatbots, are becoming an increasingly common mechanism for enterprises to provide support to customers and partners. In order to rate chatbo…

cs.CL2023

MaNtLE: Model-agnostic Natural Language Explainer

Rakesh R. Menon, Kerem Zaman, Shashank Srivastava

Understanding the internal reasoning behind the predictions of machine learning systems is increasingly vital, given their rising adoption and acceptance. While previous approaches…

cs.CL2022

What do Large Language Models Learn beyond Language?

Avinash Madasu, Shashank Srivastava

Large language models (LMs) have rapidly become a mainstay in Natural Language Processing. These models are known to acquire rich linguistic knowledge from training on large amount…

cs.CL2022

CLUES: A Benchmark for Learning Classifiers using Natural Language Explanations

Rakesh R Menon, Sayan Ghosh, Shashank Srivastava

Supervised learning has traditionally focused on inductive learning by observing labeled examples of a task. In contrast, humans have the ability to learn new concepts from languag…

cs.CL2021

Mapping Language to Programs using Multiple Reward Components with Inverse Reinforcement Learning

Sayan Ghosh, Shashank Srivastava

Mapping natural language instructions to programs that computers can process is a fundamental challenge. Existing approaches focus on likelihood-based training or using reinforceme…

cs.CL2021

Adversarial Scrubbing of Demographic Information for Text Classification

Somnath Basu Roy Chowdhury, Sayan Ghosh, Yiyuan Li +3

Contextual representations learned by language models can often encode undesirable attributes, like demographic associations of the users, while being trained for an unrelated targ…