9 papers
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery
Shahar Levy, Eliya Habba, Reshef Mintz +3
Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, traditionally obtained by manually de…
Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning
Reef Menaged, Gili Lior, Shauli Ravfogel +2
We propose agentic automata learning to evaluate the extent to which tool-calling LLM agents can uncover hidden environments through interaction. In our setup, an agent should unco…
Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation
Gili Lior, Tzviel Frostig, Gabriel Stanovsky +1
Multilingual benchmarks are central to evaluating large language models (LLMs) across languages, but they suffer from three issues: exhaustive evaluation scales linearly with the n…
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
Itay Itzhak, Eliya Habba, Gabriel Stanovsky +1
Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vibe-testing'': informal experience-based ev…
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
Eliya Habba, Noam Dahan, Gili Lior +1
Evaluating LLMs with a single prompt has proven unreliable, with small changes leading to significant performance differences. However, generating the prompt variations needed for…
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
Eliya Habba, Ofir Arviv, Itay Itzhak +5
Recent work found that LLMs are sensitive to a wide range of arbitrary prompt dimensions, including the type of delimiters, answer enumerators, instruction wording, and more. This…