activity
20242026
collaborators

6 papers

cs.LG2026

Robustness as an Emergent Property of Task Performance

Shir Ashury-Tahan, Ariel Gera, Elron Bandel +2

Robustness is often regarded as a critical future challenge for real-world applications, where stability is essential. However, as models often learn tasks in a similar order, we h…

cs.CL2026

Will it Merge? On The Causes of Model Mergeability

Adir Rahamim, Asaf Yehudai, Boaz Carmeli +3

Model merging has emerged as a promising technique for combining multiple fine-tuned models into a single multitask model without retraining. However, the factors that determine wh…

cs.CL2025

CRISP: Complex Reasoning with Interpretable Step-based Plans

Matan Vetzler, Koren Lazar, Guy Uziel +3

Recent advancements in large language models (LLMs) underscore the need for stronger reasoning capabilities to solve complex problems effectively. While Chain-of-Thought (CoT) reas…

cs.CL2025

Can Gradient Descent Simulate Prompting?

Eric Zhang, Leshem Choshen, Jacob Andreas

There are two primary ways of incorporating new information into a language model (LM): changing its prompt or changing its parameters, e.g. via fine-tuning. Parameter updates incu…

cs.CL2025

NeurIPS 2023 LLM Efficiency Fine-tuning Competition

Mark Saroufim, Yotam Perlitz, Leshem Choshen +11

Our analysis of the NeurIPS 2023 large language model (LLM) fine-tuning competition revealed the following trend: top-performing models exhibit significant overfitting on benchmark…

cs.LG2024

Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families

Felipe Maia Polo, Seamus Somerstep, Leshem Choshen +2

Scaling laws for large language models (LLMs) predict model performance based on parameters like size and training data. However, differences in training configurations and data pr…