Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models
Longwei Xu, Feng Feng, Shaojie Zhang +7
Optical Character Recognition (OCR) is increasingly regarded as a foundational capability for modern vision-language models (VLMs), enabling them not only to read text in images bu…
cs.CV2026
MASQuant: Modality-Aware Smoothing Quantization for Multimodal Large Language Models
Lulu Hu, Wenhu Xiao, Xin Chen +4
Post-training quantization (PTQ) with computational invariance for Large Language Models~(LLMs) have demonstrated remarkable advances, however, their application to Multimodal Larg…
cs.CV2024
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Pan Zhang, Xiaoyi Dong, Yuhang Cao +26
Creating AI systems that can interact with environments over long periods, similar to human cognition, has been a longstanding research goal. Recent advancements in multimodal larg…