4 papers
Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation
Luca Zhou, Sajel Shah, Emanuele Rodolà +1
Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficulty signal. The same signal drives RL with…
Tail-Shape Estimation in LLM Evaluation Is Fragile: A Protocol for Diagnosing False Positives
Luca Zhou
Recent work motivates moving large language model (LLM) evaluation from mean-based to tail-aware metrics, including conditional value-at-risk and tail-index estimates of reward-mod…
On Task Vectors and Gradients
Luca Zhou, Daniele Solombrino, Donato Crisostomi +4
Task arithmetic has emerged as a simple yet powerful technique for model merging, enabling the combination of multiple finetuned models into one. Despite its empirical success, a c…
ATM: Improving Model Merging by Alternating Tuning and Merging
Luca Zhou, Daniele Solombrino, Donato Crisostomi +3
Model merging has emerged as a cost-efficient approximation to multitask learning. Among merging strategies, task arithmetic is notable for its simplicity and effectiveness. In thi…