most citedD-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning

1 citations · 1 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

17 papers · 1 filter

cs.CV2026

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

Hao Yu, Kang Liu, Linnan Zhao +4

Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are bias…

cs.CV2026

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

Junjie Zhou, Ke Mei, Lei Li +3

Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retr…

cs.CV2026

Orca: The World is in Your Mind

Yihao Wang, Yuheng Ji, Mingyu Cao +54

We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals and exposes it through multi…

cs.CV2026

MMAgent-R: Learning to Rerank and Reject for Agentic mRAG

Tao Zhang, Ziqi Zhang, Zongyang Ma +7

Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic knowledge bases and answer rel…

cs.CV2026

Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning

Yake Wei, Yuan Wang, Fengyun Rao +2

Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with i…

cs.CV2026

Semantic-Enriched Latent Visual Reasoning

Tianrun Xu, Yue Sun, Qixun Wang +8

Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. However, existing approaches larg…