collaborators

6 papers

cs.CL2026

Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning

Reef Menaged, Gili Lior, Shauli Ravfogel +2

We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction. In our setup, an agent should unco…

cs.CL2026

Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation

Gili Lior, Tzviel Frostig, Gabriel Stanovsky +1

Multilingual benchmarks are central to evaluating large language models (LLMs) across languages, but they suffer from three issues: exhaustive evaluation scales linearly with the n…

cs.CL2026

WildIFEval: Instruction Following in the Wild

Gili Lior, Asaf Yehudai, Ariel Gera +1

Recent LLMs have shown remarkable success in following user instructions, yet handling instructions with multiple constraints remains a significant challenge. In this work, we intr…

cs.CL2026

PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation

Eliya Habba, Noam Dahan, Gili Lior +1

Evaluating LLMs with a single prompt has proven unreliable, with small changes leading to significant performance differences. However, generating the prompt variations needed for…

cs.CL2026

Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts

Gili Lior, Liron Nacchace, Gabriel Stanovsky

Humans are influenced by how information is presented, a phenomenon known as the framing effect. Prior work suggests that LLMs may also be susceptible to framing, but it has relied…

cs.CL2025

ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments

Gili Lior, Eliya Habba, Shahar Levy +2

LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations…