collaborators
Showing cs.CVShow all

22 papers · 1 filter

cs.CV2026

When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs

Congyang Ou, Ruike Song, Yang Zhou +3

Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on simil…

cs.CV2026

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models

Pengjie Wang, Linger Deng, Zujia Zhang +6

Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but typically generate each visua…

cs.CV2026

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation

Jiahao Lyu, Pei Fu, Zhenhang Li +6

In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appea…

cs.CV2026

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models

Longwei Xu, Feng Feng, Shaojie Zhang +7

Optical Character Recognition (OCR) is increasingly regarded as a foundational capability for modern vision-language models (VLMs), enabling them not only to read text in images bu…

cs.CV2026

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering

Yiduo Jia, Muzhi Zhu, Hao Zhong +7

To extend the reinforcement learning post-training paradigm to omni-modal models for concurrently bolstering video-audio understanding and collaborative reasoning, we propose OmniJ…

cs.CV2026

ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

Yiran Guan, Sifan Tu, Dingkang Liang +6

Omni-modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing omni-modal large language models (OLLM) excel…