1 paper · 1 filter
Yuxia Wang, Shimin Tao, Ning Xie +3
Despite the subjective nature of semantic textual similarity (STS) and pervasive disagreements in STS annotation, existing benchmarks have used averaged human ratings as the gold s…