From the 1 of 16 linked papers with an AI index.
2 citations · 3 across the 7 of their papers we have counts for
16 papers
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
Junjie Zhou, Ke Mei, Lei Li +3
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retr…
Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
Peixi Wu, Ke Mei, Feipeng Ma +15
The paper introduces RIME, a rewrite-driven framework that improves multimodal embeddings by jointly optimizing generation and retrieval-friendly rewriting, aligning generative and…
Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning
Yake Wei, Yuan Wang, Fengyun Rao +2
Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with i…
Semantic-Enriched Latent Visual Reasoning
Tianrun Xu, Yue Sun, Qixun Wang +8
Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches larg…
Stage-adaptive Token Selection for Efficient Omni-modal LLMs
Zijie Xin, Jie Yang, Ruixiang Zhao +4
Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequences interleaved at the window…
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
Ruixiang Zhao, Jie Yang, Zijie Xin +4
Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emerging capability of omni-moda…