Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Static Key Attention in Vision
Zizhao Hu, Xiaolin Zhou, Mohammad Rostami
The success of vision transformers is widely attributed to the expressive power of their dynamically parameterized multi-head self-attention mechanism. We examine the impact of sub…
cs.CV2023
Efficient Multimodal Diffusion Models Using Joint Data Infilling with Partially Shared U-Net
Zizhao Hu, Shaochong Jia, Mohammad Rostami
Recently, diffusion models have been used successfully to fit distributions for cross-modal data translation and multimodal data generation. However, these methods rely on extensiv…