4 citations · 8 across the 5 of their papers we have counts for
5 papers
GPT-4V(ision) as A Social Media Analysis Engine
Hanjia Lyu, Jinfa Huang, Daoan Zhang +6
Recent research has offered insights into the extraordinary capabilities of Large Multimodal Models (LMMs) in various general vision and language tasks. There is growing interest i…
Improving Scene Graph Generation with Superpixel-Based Interaction Learning
Jingyi Wang, Can Zhang, Jinfa Huang +2
Recent advances in Scene Graph Generation (SGG) typically model the relationships among entities utilizing box-level features from pre-defined detectors. We argue that an overlooke…
Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment
Peng Jin, Hao Li, Zesen Cheng +5
Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local de…
Cross-Modality Time-Variant Relation Learning for Generating Dynamic Scene Graphs
Jingyi Wang, Jinfa Huang, Can Zhang +1
Dynamic scene graphs generated from video clips could help enhance the semantic visual understanding in a wide range of challenging tasks such as environmental perception, autonomo…
Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning
Peng Jin, Jinfa Huang, Pengfei Xiong +5
Contrastive learning-based video-language representation learning approaches, e.g., CLIP, have achieved outstanding performance, which pursue semantic interaction upon pre-defined…