6 papers
Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence
Wenzhe Yin, Zehao Xiao, Pan Zhou +4
Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize…
Towards Uniformity and Alignment for Multimodal Representation Learning
Wenzhe Yin, Pan Zhou, Zehao Xiao +4
Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-…
Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-shot Semantic Segmentation
Jie Liu, Jiayi Shen, Pan Zhou +2
Generalized Few-Shot Semantic Segmentation (GFSS) aims to extend a segmentation model to novel classes with only a few annotated examples while maintaining performance on base clas…
Probabilistic Interactive 3D Segmentation with Hierarchical Neural Processes
Jie Liu, Pan Zhou, Zehao Xiao +4
Interactive 3D segmentation has emerged as a promising solution for generating accurate object masks in complex 3D scenes by incorporating user-provided clicks. However, two critic…
DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image
Qi Zhao, Zhan Ma, Pan Zhou
Recent developments in generative diffusion models have turned many dreams into realities. For video object insertion, existing methods typically require additional information, su…
CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation
Jie Liu, Pan Zhou, Yingjun Du +4
In this work, we address the cooperation problem among large language model (LLM) based embodied agents, where agents must cooperate to achieve a common goal. Previous methods ofte…