5 papers · 1 filter
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
Md Nayem Uddin, Kumar Shubham, Eduardo Blanco +2
Personalized agents that interact with users over long periods must maintain persistent memory across sessions and update it as circumstances change. However, existing benchmarks p…
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
Md Nayem Uddin, Amir Saeidi, Divij Handa +5
This paper introduces UnSeenTimeQA, a novel data contamination-free time-sensitive question-answering (TSQA) benchmark. It differs from existing TSQA benchmarks by avoiding web-sea…
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
Aswin RRV, Nemika Tyagi, Md Nayem Uddin +2
This study explores the sycophantic tendencies of Large Language Models (LLMs), where these models tend to provide answers that match what users want to hear, even if they are not…
Asking and Answering Questions to Extract Event-Argument Structures
Md Nayem Uddin, Enfa Rose George, Eduardo Blanco +1
This paper presents a question-answering approach to extract document-level event-argument structures. We automatically ask and answer questions for each argument type an event may…
Generating Uncontextualized and Contextualized Questions for Document-Level Event Argument Extraction
Md Nayem Uddin, Enfa Rose George, Eduardo Blanco +1
This paper presents multiple question generation strategies for document-level event argument extraction. These strategies do not require human involvement and result in uncontextu…