most citedVideo-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning

3 citations · 5 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2023

Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment

Peng Jin, Hao Li, Zesen Cheng +5

Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local de…

cs.CV2023

TG-VQA: Ternary Game of Video Question Answering

Hao Li, Peng Jin, Zesen Cheng +5

Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions…

cs.CV20233 cited

Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning

Peng Jin, Jinfa Huang, Pengfei Xiong +5

Contrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined…

cs.CV2023

Multi-granularity Interaction Simulation for Unsupervised Interactive Segmentation

Kehan Li, Yian Zhao, Zhennan Wang +6

Interactive segmentation enables users to segment as needed by providing cues of objects, which introduces human-computer interaction for many fields, such as image editing and med…

cs.CV20232 cited

Parallel Vertex Diffusion for Unified Visual Grounding

Zesen Cheng, Kehan Li, Peng Jin +4

Unified visual grounding pursues a simple and generic technical route to leverage multi-task data with less task-specific design. The most advanced methods typically present boxes…