33 citations · 45 across the 6 of their papers we have counts for
6 papers
Instruction-Following Evaluation for Large Language Models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra +5
One core capability of Large Language Models (LLMs) is to follow natural language instructions. However, the evaluation of such abilities is not standardized: Human evaluations are…
InstructExcel: A Benchmark for Natural Language Instruction in Excel
Justin Payan, Swaroop Mishra, Mukul Singh +7
With the evolution of Large Language Models (LLMs) we can solve increasingly more complex NLP tasks across various domains, including spreadsheets. This work investigates whether L…
How FaR Are Large Language Models From Agents with Theory-of-Mind?
Pei Zhou, Aman Madaan, Srividya Pranavi Potharaju +9
"Thinking is for Doing." Humans can infer other people's mental states from observations--an ability called Theory-of-Mind (ToM)--and subsequently act pragmatically on those infere…
Instruction Tuned Models are Quick Learners
Himanshu Gupta, Saurabh Arjun Sawant, Swaroop Mishra +4
Instruction tuning of language models has demonstrated the ability to enhance model generalization to unseen tasks via in-context learning using a few examples. However, typical su…
Real-Time Visual Feedback to Guide Benchmark Creation: A Human-and-Metric-in-the-Loop Workflow
Anjana Arunkumar, Swaroop Mishra, Bhavdeep Sachdeva +2
Recent research has shown that language models exploit `artifacts' in benchmarks to solve tasks, rather than truly learning them, leading to inflated model performance. In pursuit…
BioTABQA: Instruction Learning for Biomedical Table Question Answering
Man Luo, Sharad Saxena, Swaroop Mishra +2
Table Question Answering (TQA) is an important but under-explored task. Most of the existing QA datasets are in unstructured text format and only few of them use tables as the cont…