3 papers
cs.CL2026
STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation
Bhiman Kumar Baghel, Anna Chrabaszcz, Tessa Warren +3
Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language processing. To isolate plausi…
cs.CL2026
CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity
Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu +9
Creativity is often seen as a hallmark of human intelligence. While large language models(LLMs) are increasingly perceived as generating creative text, there is still no cross-doma…
cs.CL2025
Resolving UnderEdit & OverEdit with Iterative & Neighbor-Assisted Model Editing
Bhiman Kumar Baghel, Emma Jordan, Zheyuan Ryan Shi +1
Large Language Models (LLMs) are widely deployed in downstream tasks, but keeping their knowledge up-to-date via retraining or fine-tuning is often computationally expensive. Model…