1 paper
Kathy Garcia, Leyla Isik
Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social information in dynamic scenes. For examp…