10 papers
Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding
Shiyu Li, Zhiyuan Hu, Yifan Wang +3
The paper presents Conan-embedding-v3, a framework that trains modality‑specific specialist models, fuses them into a single dense backbone, and then recovers the modality projecto…
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
Peiming Li, Yifan Wang, Xiaotian Zhang +4
Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step perception and logical deduct…
Focus When Necessary: Adaptive Routing and Collaborative Grounding for Training-Free Visual Grounding
Yifan Wang, Peiming Li, Shiyu Li +5
While Multimodal Large Language Models (MLLMs) excel in cross-modal reasoning, they often struggle to perceive fine-grained details in complex high-resolution images. Recent traini…
Aligning Deep Implicit Preferences by Learning to Reason Defensively
Peiming Li, Zhiyuan Hu, Yang Tang +2
Personalized alignment is crucial for enabling Large Language Models (LLMs) to engage effectively in user-centric interactions. However, current methods face a dual challenge: they…
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
Yifan Wang, Shiyu Li, Peiming Li +3
Chain-of-Thought (CoT) prompting has achieved remarkable success in unlocking the reasoning capabilities of Large Language Models (LLMs). Although CoT prompting enhances reasoning,…
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
Ziyi Wang, Peiming Li, Xinshun Wang +3
Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet cannot process structured, non-visual data such as human skeletons. Existing methods either c…