activity
20202024
most citedExtending Multi-modal Contrastive Representations

2 citations · 2 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CV2024

FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion

Zehan Wang, Ziang Zhang, Xize Cheng +8

Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting probl…

cs.CV2023

Multi-Modal Domain Adaptation Across Video Scenes for Temporal Video Grounding

Haifeng Huang, Yang Zhao, Zehan Wang +2

Temporal Video Grounding (TVG) aims to localize the temporal boundary of a specific segment in an untrimmed video based on a given language query. Since datasets in this domain are…

cs.CV2023

Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Haifeng Huang, Yilun Chen, Zehan Wang +8

Recent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding. However, previous methods exhibit deficiencies in genera…

cs.CV20232 cited

Extending Multi-modal Contrastive Representations

Zehan Wang, Ziang Zhang, Luping Liu +4

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high d…

cs.CV2020

Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences

Zhu Zhang, Zhou Zhao, Yang Zhao +3

In this paper, we consider a novel task, Spatio-Temporal Video Grounding for Multi-Form Sentences (STVG). Given an untrimmed video and a declarative/interrogative sentence depictin…