7 papers · 1 filter
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
Junjie Zhou, Ke Mei, Lei Li +3
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retr…
AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer
Junqiu Yu, Pandeng Li, Yikai Wang +11
Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be organized to facilitate denoisi…
Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding
Jiapeng Li, Yong Li, Junjie Zhou +2
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region p…
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
Yujie Wei, Yujin Han, Zhekai Chen +20
Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However, evaluating such frontier mo…
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
Peixi Wu, Ke Mei, Feipeng Ma +15
Multimodal Large Language Models (MLLMs) have emerged as a promising foundation for universal multimodal embeddings. Recent studies have shown that reasoning-driven generative mult…
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
Kaixun Jiang, Yuzheng Wang, Junjie Zhou +6
We introduce GenAgent, unifying visual understanding and generation through an agentic multimodal model. Unlike unified models that face expensive training costs and understanding-…