Score-Only Distillation for Compact Dense Retrieval
arXiv:2607.11465
The paper proposes a method to compress large dense retrieval models into much smaller ones by distilling only the teacher's score vectors, using a row‑centered score‑vector objective and a uniform all‑pairs PairMSE loss, achieving up to 50% of the performance gap while greatly speeding up encoding.
Abstract
Large embedding models improve retrieval quality, but serving large encoders online is expensive. We study whether a compact retriever can learn teacher ranking behavior from score vectors without access to teacher hidden states. The student trains on rows built from ground-truth positives and negative candidates produced by our data generation pipeline; we evaluate student-teacher hard-negative mining separately as an extension. We use a row-centered score-vector objective, a memory-efficient implementation of uniform all-pairs PairMSE loss. On a fixed eight-task evaluation panel, our distillation protocol recovers up to 50\% of the base-to-teacher gap. The distilled 0.6B student is 4.7 faster for query encoding and 9.7 faster for document encoding than sequential online teacher fusion. External-transfer performance after distillation remains mixed, so our evidence supports compression of teacher rankings under matched retrieval protocols.