62 citations · 268 across the 69 of their papers we have counts for
57 papers · 1 filter
When Linguistic and Internal Confidence Diverge in Large Language Models
Hefan Zhang, Bingquan Zhang, Ming Cheng +3
Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study…
Improving Information Extraction with Learned Queries
Omar Sharif, Soroush Vosoughi, Nikhil Singh
When information extraction fails, a natural instinct is to improve the model doing it: for example, by scaling it up or refining its reasoning. In this paper, we show that another…
UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory
Peijun Qing, Fobo Shi, Soroush Vosoughi
Long-term memory is increasingly important for conversational agents, yet existing benchmarks primarily measure memory through pointwise factual recall: whether a system can recove…
When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation
Yahan Zheng, John Guerrerio, Soroush Vosoughi +1
Preprocessing-based methods for stereotype mitigation, such as pre-/post-training on debiased corpora, are widely used in NLP. While these approaches reduce measurable stereotypes…
Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration
Weicheng Ma, John Guerrerio, Soroush Vosoughi
Research on stereotypes in large language models (LLMs) has largely focused on English-speaking contexts, due to the lack of datasets in other languages and the high cost of manual…
Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agents
Yuxin Wang, Paul Thomas, Zhiwei Yu +5
Prior research on memory mechanism in RAG-based conversational system has emphasized how memory is stored and retrieved. However, far less is known about how memories with differen…