1k citations · 1.3k across the 10 of their papers we have counts for
Showing 2023 · cs.CVShow all
2 papers · 2 filters
cs.CV2023
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
Rui Sun, Zhecan Wang, Haoxuan You +3
Vision-language tasks, such as VQA, SNLI-VE, and VCR are challenging because they require the model's reasoning ability to understand the semantics of the visual world and natural…
cs.CV2023
Streaming Video Model
Yucheng Zhao, Chong Luo, Chuanxin Tang +3
Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recog…