Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Rethinking Causal Mask Attention for Vision-Language Inference
Xiaohuan Pei, Tao Huang, YanXiang Ma +1
Causal attention has become a foundational mechanism in autoregressive vision-language models (VLMs), unifying textual and visual inputs under a single generative framework. Howeve…
cs.CV2025
Learning Mask Invariant Mutual Information for Masked Image Modeling
Tao Huang, Yanxiang Ma, Shan You +1
Masked autoencoders (MAEs) represent a prominent self-supervised learning paradigm in computer vision. Despite their empirical success, the underlying mechanisms of MAEs remain ins…