Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
In-Context Collapse in Vision-Language Models and How to Mitigate it?
Mohammad Rostami
Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely assumed to improve as more demon…
cs.CV2024
Static Key Attention in Vision
Zizhao Hu, Xiaolin Zhou, Mohammad Rostami
The success of vision transformers is widely attributed to the expressive power of their dynamically parameterized multi-head self-attention mechanism. We examine the impact of sub…
cs.CV2024
Lateralization MLP: A Simple Brain-inspired Architecture for Diffusion
Zizhao Hu, Mohammad Rostami
The Transformer architecture has dominated machine learning in a wide range of tasks. The specific characteristic of this architecture is an expensive scaled dot-product attention…