1 citations · 1 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation
Luca Zhou, Sajel Shah, Emanuele Rodolà +1
Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficulty signal. The same signal drives RL with…
cs.LG2025
Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
Davide Marincione, Donato Crisostomi, Roberto Dessi +2
Foundation models capable of generalizing across species and tasks represent a promising new frontier in bioacoustics, with NatureLM being one of the most prominent examples. While…