6 papers
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
Ariel Gera, Shir Ashury-Tahan, Gal Bloch +2
We explore the effectiveness of an LLM-guided query refinement paradigm for extending the usability of embedding models to challenging zero-shot search and classification tasks. Ou…
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
Shir Ashury-Tahan, Yifan Mai, Elron Bandel +2
Large Language Models (LLM) benchmarks tell us when models fail, but not why they fail. A wrong answer on a reasoning dataset may stem from formatting issues, calculation errors, o…
Robustness as an Emergent Property of Task Performance
Shir Ashury-Tahan, Ariel Gera, Elron Bandel +2
Robustness is widely viewed as a key challenge for real-world applications. However, because current research focuses only on difficult tasks, it partially captures real-world read…
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
Shir Ashury-Tahan, Yifan Mai, Rajmohan C +8
Despite its real-world significance, model performance on tabular data remains underexplored, leaving uncertainty about which model to rely on and which prompt configuration to ado…
Data-driven Coreference-based Ontology Building
Shir Ashury-Tahan, Amir David Nissan Cohen, Nadav Cohen +2
While coreference resolution is traditionally used as a component in individual document understanding, in this work we take a more global view and explore what can we learn about…
Label-Efficient Model Selection for Text Generation
Shir Ashury-Tahan, Ariel Gera, Benjamin Sznajder +3
Model selection for a given target task can be costly, as it may entail extensive annotation of the quality of outputs of different models. We introduce DiffUse, an efficient metho…