collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation

Jiahao Lyu, Pei Fu, Zhenhang Li +6

In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appea…

cs.CV2026

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models

Longwei Xu, Feng Feng, Shaojie Zhang +7

Optical Character Recognition (OCR) is increasingly regarded as a foundational capability for modern vision-language models (VLMs), enabling them not only to read text in images bu…

cs.CV2026

PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues

Yukun Qi, Pei Fu, Hang Li +5

Vision-Language Models (VLMs) have achieved remarkable progress on a wide range of challenging multimodal understanding and reasoning tasks. However, existing reasoning paradigms,…

cs.CV2025

Xiaomi MiMo-VL-Miloco Technical Report

Jiaze Li, Jingyang Chen, Yuxun Qu +9

We open-source MiMo-VL-Miloco-7B and its quantized variant MiMo-VL-Miloco-7B-GGUF, a pair of home-centric vision-language models that achieve strong performance on both home-scenar…

cs.CV2025

BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent

Shaojie Zhang, Ruoceng Zhang, Pei Fu +8

In the field of AI-driven human-GUI interaction automation, while rapid advances in multimodal large language models and reinforcement fine-tuning techniques have yielded remarkabl…