1 paper
Tz-Ying Wu, Kyle Min, Subarna Tripathi +1
Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) bas…