Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations
Yindong Wang, Martin PreiÃ, Margarita Bugueño +4
The mechanisms underlying scientific confabulation in Large Language Models (LLMs) remain poorly understood. We introduce ReFACT (Reddit False And Correct Texts), a benchmark of 1,…
cs.CL2024
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit's Showerthoughts
Tolga Buz, Benjamin Frost, Nikola Genchev +3
Recent Large Language Models (LLMs) have shown the ability to generate content that is difficult or impossible to distinguish from human writing. We investigate the ability of diff…