1 citations · 1 across the 15 of their papers we have counts for
17 papers · 1 filter
WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing
Hao Yu, Kang Liu, Linnan Zhao +4
Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are bias…
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
Junjie Zhou, Ke Mei, Lei Li +3
Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retr…
Orca: The World is in Your Mind
Yihao Wang, Yuheng Ji, Mingyu Cao +54
We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals and exposes it through multi…
MMAgent-R: Learning to Rerank and Reject for Agentic mRAG
Tao Zhang, Ziqi Zhang, Zongyang Ma +7
Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic knowledge bases and answer rel…
Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning
Yake Wei, Yuan Wang, Fengyun Rao +2
Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with i…
Semantic-Enriched Latent Visual Reasoning
Tianrun Xu, Yue Sun, Qixun Wang +8
Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches larg…