activity
20162024
most citedNeedle in the Haystack for Memory Based Large Language Models

5 citations · 8 across the 10 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2024

NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation

Karan Wanchoo, Xiaoye Zuo, Hannah Gonzalez +5

We present NAVCON, a large-scale annotated Vision-Language Navigation (VLN) corpus built on top of two popular datasets (R2R and RxR). The paper introduces four core, cognitively m…

cs.CL2024

Benchmarking LLM Guardrails in Handling Multilingual Toxicity

Yahan Yang, Soham Dan, Dan Roth +1

With the ubiquity of Large Language Models (LLMs), guardrails have become crucial to detect and defend against toxic content. However, with the increasing pervasiveness of LLMs in…

cs.CL2024

Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models

Amey Hengle, Prasoon Bajpai, Soham Dan +1

While recent large language models (LLMs) demonstrate remarkable abilities in responding to queries in diverse languages, their ability to handle long multilingual contexts is unex…

cs.CL20245 cited

Needle in the Haystack for Memory Based Large Language Models

Elliot Nelson, Georgios Kollias, Payel Das +2

Current large language models (LLMs) often perform poorly on simple fact retrieval tasks. Here we investigate if coupling a dynamically adaptable external memory to a LLM can allev…

cs.CL20243 cited

CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design

Nafis Neehal, Bowen Wang, Shayom Debopadhaya +4

CTBench is introduced as a benchmark to assess language models (LMs) in aiding clinical study design. Given study-specific metadata, CTBench evaluates AI models' ability to determi…

cs.CL2024

On the Effects of Fine-tuning Language Models for Text-Based Reinforcement Learning

Mauricio Gruppi, Soham Dan, Keerthiram Murugesan +1

Text-based reinforcement learning involves an agent interacting with a fictional environment using observed text and admissible actions in natural language to complete a task. Prev…