3 papers
cs.CL2026
From Benchmarks to Skills: Low-Rank Factors for LLM Evaluation
Aviya Maimon, Amir DN Cohen, Gal Vishne +2
Current evaluations of large language models (LLMs) rely heavily on a growing collection of benchmarks and on aggregate benchmark scores, yet it remains unclear what this compariso…
cs.CL2025
Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
Omer Goldman, Alon Jacovi, Aviv Slobodkin +3
Improvements in language models' capabilities have pushed their applications towards longer contexts, making long-context evaluation and development an active research area. Howeve…
cs.CV2024
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
Nitzan Bitton-Guetta, Aviv Slobodkin, Aviya Maimon +6
Imagine observing someone scratching their arm; to understand why, additional context would be necessary. However, spotting a mosquito nearby would immediately offer a likely expla…