1 paper · 1 filter
Adrian Grassi
Static benchmarks measure a model frozen at training time. Real systems face distribution shift: new categories, paraphrased queries, drift: and must recover online via user correc…