3 citations · 3 across the 4 of their papers we have counts for
1 paper · 1 filter
Sihyun Yu, Nanye Ma, Pinzhi Huang +8
Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacity for visual state tracking…