6 papers
Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning
Reef Menaged, Gili Lior, Shauli Ravfogel +2
We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction. In our setup, an agent should unco…
Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation
Gili Lior, Tzviel Frostig, Gabriel Stanovsky +1
Multilingual benchmarks are central to evaluating large language models (LLMs) across languages, but they suffer from three issues: exhaustive evaluation scales linearly with the n…
WildIFEval: Instruction Following in the Wild
Gili Lior, Asaf Yehudai, Ariel Gera +1
Recent LLMs have shown remarkable success in following user instructions, yet handling instructions with multiple constraints remains a significant challenge. In this work, we intr…
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
Eliya Habba, Noam Dahan, Gili Lior +1
Evaluating LLMs with a single prompt has proven unreliable, with small changes leading to significant performance differences. However, generating the prompt variations needed for…
Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts
Gili Lior, Liron Nacchace, Gabriel Stanovsky
Humans are influenced by how information is presented, a phenomenon known as the framing effect. Prior work suggests that LLMs may also be susceptible to framing, but it has relied…
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
Gili Lior, Eliya Habba, Shahar Levy +2
LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations…