1 paper
Jiapeng Shi, Junke Wang, Zuyao You +2
This paper presents VideoLoom, a unified Video Large Language Model (Video LLM) for joint spatial-temporal understanding. To facilitate the development of fine-grained spatial and…