1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2025
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
Elon Ezra, Ariel Weizman, Amos Azaria
Large language models (LLMs) are commonly evaluated on tasks that test their knowledge or reasoning abilities. In this paper, we explore a different type of evaluation: whether an…
cs.CL2024★ 1 cited
Fool Me, Fool Me: User Attitudes Toward LLM Falsehoods
Diana Bar-Or Nirman, Ariel Weizman, Amos Azaria
While Large Language Models (LLMs) have become central tools in various fields, they often provide inaccurate or false information. This study examines user preferences regarding f…