13 citations · 45 across the 24 of their papers we have counts for
9 papers · 1 filter
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Jared Moore, Andrea Mock, Yifan Mai +9
Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including "delusional spirals" in which conce…
Characterizing Delusional Spirals through Human-LLM Chat Logs
Jared Moore, Ashish Mehta, William Agnew +11
As large language models (LLMs) have proliferated, disturbing anecdotal reports of negative psychological effects, such as delusions, self-harm, and ``AI psychosis,'' have emerged…
OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasets
Jiyuan Shen, Peiyue Yuan, Atin Ghosh +2
Multimodal Large Language Models (MLLMs) enhance the potential of natural language processing. However, their actual impact on document information extraction remains unclear. In p…
Structured Prompts Improve Evaluation of Language Models
Asad Aali, Muhammad Ahmed Mohsin, Vasiliki Bikia +15
As language models (LMs) are increasingly adopted across domains, high-quality benchmarking frameworks are essential for guiding deployment decisions. In practice, however, framewo…
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes +78
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
Shir Ashury-Tahan, Yifan Mai, Rajmohan C +8
Despite its real-world significance, model performance on tabular data remains underexplored, leaving uncertainty about which model to rely on and which prompt configuration to ado…