1 paper
Sharon S. Musa, Fereshteh Forghani, Harrish Thasarathan +3
Self-supervised video foundation models learn rich spatiotemporal representations, yet it remains unclear what visual concepts these representations encode, where they emerge acros…