activity
20232026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

Ziang Wu, Peng Jin, Qishen Yin +4

Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this token mix. We call the stan…

cs.CV2024

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning

Peng Jin, Hao Li, Li Yuan +2

Multimodal representation learning, with contrastive learning, plays an important role in the artificial intelligence domain. As an important subfield, video-language representatio…

cs.CV2024

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Guowei Xu, Peng Jin, Ziang Wu +4

Large language models have demonstrated substantial advancements in reasoning capabilities. However, current Vision-Language Models (VLMs) often struggle to perform systematic and…

cs.CV2024

Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation

Peng Jin, Hao Li, Zesen Cheng +6

Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global mot…

cs.CV2024

Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation

Zesen Cheng, Kehan Li, Hao Li +5

Temporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video…

cs.CV2023

FreestyleRet: Retrieving Images from Style-Diversified Queries

Hao Li, Curise Jia, Peng Jin +5

Image Retrieval aims to retrieve corresponding images based on a given query. In application scenarios, users intend to express their retrieval intent through various query styles.…