2 citations · 3 across the 4 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
Autoregressive Pretraining with Mamba in Vision
Sucheng Ren, Xianhang Li, Haoqin Tu +9
The vision community has started to build with the recently developed state space model, Mamba, as the new backbone for a range of tasks. This paper shows that Mamba's visual capab…
cs.CV2023
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
Mingfei Han, Linjie Yang, Xiaojun Chang +2
A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to…
cs.CV2023★ 2 cited
The Devil is in the Details: A Deep Dive into the Rabbit Hole of Data Filtering
Haichao Yu, Yu Tian, Sateesh Kumar +2
The quality of pre-training data plays a critical role in the performance of foundation models. Popular foundation models often design their own recipe for data filtering, which ma…