2 papers
cs.CV2024
DRDM: A Disentangled Representations Diffusion Model for Synthesizing Realistic Person Images
Enbo Huang, Yuan Zhang, Faliang Huang +2
Person image synthesis with controllable body poses and appearances is an essential task owing to the practical needs in the context of virtual try-on, image editing and video prod…
cs.CV2024
VisionGRU: A Linear-Complexity RNN Model for Efficient Image Analysis
Shicheng Yin, Kaixuan Yin, Weixing Chen +2
Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) are two dominant models for image analysis. While CNNs excel at extracting multi-scale features and ViTs effecti…