2 papers
cs.CL2025
MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
Tomer Wolfson, Harsh Trivedi, Mor Geva +5
Automated agents, powered by Large language models (LLMs), are emerging as the go-to tool for querying information. However, evaluation benchmarks for LLM agents rarely feature nat…
cs.CL2025
HeQ: a Large and Diverse Hebrew Reading Comprehension Benchmark
Amir DN Cohen, Hilla Merhav, Yoav Goldberg +1
Current benchmarks for Hebrew Natural Language Processing (NLP) focus mainly on morpho-syntactic tasks, neglecting the semantic dimension of language understanding. To bridge this…