5 citations · 11 across the 7 of their papers we have counts for
7 papers · 1 filter
SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and Synopses
Chaolei Tan, Zihang Lin, Junfu Pu +7
Video grounding is a fundamental problem in multimodal content understanding, aiming to localize specific natural language queries in an untrimmed video. However, current video gro…
PTVD: A Large-Scale Plot-Oriented Multimodal Dataset Based on Television Dramas
Chen Li, Xutan Peng, Teng Wang +5
Art forms such as movies and television (TV) dramas are reflections of the real world, which have attracted much attention from the multimodal learning community recently. However,…
Knowledge-augmented Few-shot Visual Relation Detection
Tianyu Yu, Yangning Li, Jiaoyan Chen +8
Visual Relation Detection (VRD) aims to detect relationships between objects for image understanding. Most existing VRD methods rely on thousands of training samples of each relati…
Darwinian Model Upgrades: Model Evolving with Selective Compatibility
Binjie Zhang, Shupeng Su, Yixiao Ge +5
The traditional model upgrading paradigm for retrieval requires recomputing all gallery embeddings before deploying the new model (dubbed as "backfilling"), which is quite expensiv…
Privacy-Preserving Model Upgrades with Bidirectional Compatible Training in Image Retrieval
Shupeng Su, Binjie Zhang, Yixiao Ge +4
The task of privacy-preserving model upgrades in image retrieval desires to reap the benefits of rapidly evolving new models without accessing the raw gallery images. A pioneering…
Towards Universal Backward-Compatible Representation Learning
Binjie Zhang, Yixiao Ge, Yantao Shen +6
Conventional model upgrades for visual search systems require offline refresh of gallery features by feeding gallery images into new models (dubbed as "backfill"), which is time-co…