activity
20202026
most citedBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

565 citations · 725 across the 92 of their papers we have counts for

collaborators
Showing 2024 · cs.CLShow all

10 papers · 2 filters

cs.CL2024

ToW: Thoughts of Words Improve Reasoning in Large Language Models

Zhikun Xu, Ming Shen, Jacob Dineen +6

We introduce thoughts of words (ToW), a novel training-time data-augmentation method for next-word prediction. ToW views next-word prediction as a core reasoning task and injects f…

cs.CL2024

Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?

Nemika Tyagi, Mihir Parmar, Mohith Kulkarni +5

Solving grid puzzles involves a significant amount of logical reasoning. Hence, it is a good domain to evaluate the reasoning capability of a model which can then guide us to impro…

cs.CL2024★ 1 cited

Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation

Neeraj Varshney, Satyam Raj, Venkatesh Mishra +4

Large Language Models (LLMs) have achieved remarkable performance across a wide variety of natural language tasks. However, they have been shown to suffer from a critical limitatio…

cs.CL2024

Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models

Nisarg Patel, Mohith Kulkarni, Mihir Parmar +4

As Large Language Models (LLMs) continue to exhibit remarkable performance in natural language understanding tasks, there is a crucial need to measure their ability for human-like…

cs.CL2024

Cutting Through the Noise: Boosting LLM Performance on Math Word Problems

Ujjwala Anantheswaran, Himanshu Gupta, Kevin Scaria +3

Large Language Models (LLMs) excel at various tasks, including solving math word problems (MWPs), but struggle with real-world problems containing irrelevant information. To addres…

cs.CL2024★ 2 cited

Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies

Aswin RRV, Nemika Tyagi, Md Nayem Uddin +2

This study explores the sycophantic tendencies of Large Language Models (LLMs), where these models tend to provide answers that match what users want to hear, even if they are not…