Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
NoveltyBench: Evaluating Language Models for Humanlike Diversity
Yiming Zhang, Harshita Diddee, Susan Holm +5
Language models have demonstrated remarkable capabilities on standard benchmarks, yet they struggle increasingly from mode collapse, the inability to generate diverse and novel out…
cs.CL2025
Are Triggers Needed for Document-Level Event Extraction?
Shaden Shaar, Wayne Chen, Maitreyi Chatterjee +3
Most existing work on event extraction has focused on sentence-level texts and presumes the identification of a trigger-span -- a word or phrase in the input that evokes the occurr…
cs.CL2024
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering
Ruosen Li, Ruochen Li, Barry Wang +1
To evaluate Large Language Models (LLMs) for question answering (QA), traditional methods typically focus on assessing single-turn responses to given questions. However, this appro…