1 paper
Nikolai Warner, Cameron Ethan Taylor, Irfan Essa +1
Text-motion retrieval systems learn shared embedding spaces from motion-caption pairs via contrastive objectives. However, each caption is not a deterministic label but a sample fr…