most citedThe Mighty ToRR: A Benchmark for Table Reasoning and Robustness

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CL2026

Task-Adaptive Embedding Refinement via Test-time LLM Guidance

Ariel Gera, Shir Ashury-Tahan, Gal Bloch +2

We explore the effectiveness of an LLM-guided query refinement paradigm for extending the usability of embedding models to challenging zero-shot search and classification tasks. Ou…

cs.AI2026

ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models

Shir Ashury-Tahan, Yifan Mai, Elron Bandel +2

Large Language Models (LLM) benchmarks tell us when models fail, but not why they fail. A wrong answer on a reasoning dataset may stem from formatting issues, calculation errors, o…

cs.CL20261 cited

The Mighty ToRR: A Benchmark for Table Reasoning and Robustness

Shir Ashury-Tahan, Yifan Mai, Rajmohan C +8

Despite its real-world significance, model performance on tabular data remains underexplored, leaving uncertainty about which model to rely on and which prompt configuration to ado…

cs.LG2026

Robustness as an Emergent Property of Task Performance

Shir Ashury-Tahan, Ariel Gera, Elron Bandel +2

Robustness is often regarded as a critical future challenge for real-world applications, where stability is essential. However, as models often learn tasks in a similar order, we h…

cs.CL2024

Data-driven Coreference-based Ontology Building

Shir Ashury-Tahan, Amir David Nissan Cohen, Nadav Cohen +2

While coreference resolution is traditionally used as a component in individual document understanding, in this work we take a more global view and explore what can we learn about…

cs.CL2024

Label-Efficient Model Selection for Text Generation

Shir Ashury-Tahan, Ariel Gera, Benjamin Sznajder +3

Model selection for a given target task can be costly, as it may entail extensive annotation of the quality of outputs of different models. We introduce DiffUse, an efficient metho…