collaborators

9 papers

cs.CL2026

ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery

Shahar Levy, Eliya Habba, Reshef Mintz +3

Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually de…

cs.CL2026

Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning

Reef Menaged, Gili Lior, Shauli Ravfogel +2

We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction. In our setup, an agent should unco…

cs.CL2026

Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation

Gili Lior, Tzviel Frostig, Gabriel Stanovsky +1

Multilingual benchmarks are central to evaluating large language models (LLMs) across languages, but they suffer from three issues: exhaustive evaluation scales linearly with the n…

cs.CL2026

From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs

Itay Itzhak, Eliya Habba, Gabriel Stanovsky +1

Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vibe-testing'': informal experience-based ev…

cs.CL2026

PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation

Eliya Habba, Noam Dahan, Gili Lior +1

Evaluating LLMs with a single prompt has proven unreliable, with small changes leading to significant performance differences. However, generating the prompt variations needed for…

cs.CL2026

DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation

Eliya Habba, Ofir Arviv, Itay Itzhak +5

Recent work found that LLMs are sensitive to a wide range of arbitrary prompt dimensions, including the type of delimiters, answer enumerators, instruction wording, and more. This…