1 paper · 1 filter
Ruibo Ming, Lei Sun, Deheng Zhang +8
Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most…