479 citations · 1k across the 19 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024★ 1 cited
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
Michael S. Ryoo, Honglu Zhou, Shrikant Kendre +9
We present xGen-MM-Vid (BLIP-3-Video): a multimodal language model for videos, particularly designed to efficiently capture temporal information over multiple frames. BLIP-3-Video…
cs.CV2021★ 3 cited
Value Retrieval with Arbitrary Queries for Form-like Documents
Mingfei Gao, Le Xue, Chetan Ramaiah +3
We propose value retrieval with arbitrary queries for form-like documents to reduce human effort of processing forms. Unlike previous methods that only address a fixed set of field…
cs.CV2014★ 2 cited
Compositional Structure Learning for Action Understanding
Ran Xu, Gang Chen, Caiming Xiong +2
The focus of the action understanding literature has predominately been classification, how- ever, there are many applications demanding richer action understanding such as mobile…