1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
Chenyu Zhou, Mengdan Zhang, Peixian Chen +5
The swift progress of Multi-modal Large Models (MLLMs) has showcased their impressive ability to tackle tasks blending vision and language. Yet, most current models and benchmarks…
cs.CV2023
D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance Annotation
Hanjun Li, Xiujun Shu, Sunan He +5
Temporal sentence grounding (TSG) aims to locate a specific moment from an untrimmed video with a given natural language query. Recently, weakly supervised methods still have a lar…
cs.CV2023
Coarse-to-Fine: Learning Compact Discriminative Representation for Single-Stage Image Retrieval
Yunquan Zhu, Xinkai Gao, Bo Ke +2
Image retrieval targets to find images from a database that are visually similar to the query image. Two-stage methods following retrieve-and-rerank paradigm have achieved excellen…