activity
20172025
most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 57 across the 14 of their papers we have counts for

collaborators
Showing cs.CLShow all

22 papers · 1 filter

cs.CL2025

LM Agents for Coordinating Multi-User Information Gathering

Harsh Jhamtani, Jacob Andreas, Benjamin Van Durme

This paper introduces PeopleJoin, a benchmark for evaluating LM-mediated collaborative problem solving. Given a user request, PeopleJoin agents must identify teammates who might be…

cs.CL2024

Towards Robust Evaluation of Unlearning in LLMs via Data Transformations

Abhinav Joshi, Shaswati Saha, Divyaksh Shukla +4

Large Language Models (LLMs) have shown to be a great success in a wide range of applications ranging from regular NLP-based use cases to AI agents. LLMs have been trained on a vas…

cs.CL2024

Steering Large Language Models between Code Execution and Textual Reasoning

Yongchao Chen, Harsh Jhamtani, Srinagesh Sharma +2

While a lot of recent research focuses on enhancing the textual reasoning capabilities of Large Language Models (LLMs) by optimizing the multi-agent framework or reasoning chains,…

cs.CL2024

Learning to Retrieve Iteratively for In-Context Learning

Yunmo Chen, Tongfei Chen, Harsh Jhamtani +4

We introduce iterative retrieval, a novel framework that empowers retrievers to make iterative decisions through policy optimization. Finding an optimal portfolio of retrieved item…

cs.CL2023

Interpreting User Requests in the Context of Natural Language Standing Instructions

Nikita Moghe, Patrick Xia, Jacob Andreas +3

Users of natural language interfaces, generally powered by Large Language Models (LLMs),often must repeat their preferences each time they make a similar request. We describe an ap…

cs.CL20222 cited

PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification data for Learning Enhanced generation

Sedrick Scott Keh, Kevin Lu, Varun Gangal +4

A personification is a figure of speech that endows inanimate entities with properties and actions typically seen as requiring animacy. In this paper, we explore the task of person…