2 papers
cs.CV2026
Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective
Qiyan Zhao, Xiaofeng Zhang, Shuochen Chang +7
Recent diffusion-based Multimodal Large Language Models (dMLLMs) suffer from high inference latency and therefore rely on caching techniques to accelerate decoding. However, the ap…
cs.CV2025
TextMamba: Scene Text Detector with Mamba
Qiyan Zhao, Yue Yan, Da-Han Wang
In scene text detection, Transformer-based methods have addressed the global feature extraction limitations inherent in traditional convolution neural network-based methods. Howeve…