15 citations · 22 across the 13 of their papers we have counts for
11 papers · 1 filter
Enhanced Multi-Scale Cross-Attention for Person Image Generation
Hao Tang, Ling Shao, Nicu Sebe +1
In this paper, we propose a novel cross-attention-based generative adversarial network (GAN) for the challenging person image generation task. Cross-attention is a novel and intuit…
3D Part Segmentation via Geometric Aggregation of 2D Visual Features
Marco Garosi, Riccardo Tedoldi, Davide Boscaini +3
Supervised 3D part segmentation models are tailored for a fixed set of objects and parts, limiting their transferability to open-set, real-world scenarios. Recent works have explor…
Optimizing Resource Consumption in Diffusion Models through Hallucination Early Detection
Federico Betti, Lorenzo Baraldi, Rita Cucchiara +1
Diffusion models have significantly advanced generative AI, but they encounter difficulties when generating complex combinations of multiple objects. As the final result heavily de…
3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance
Xiaoxu Xu, Yitian Yuan, Jinlong Li +6
In this paper, we propose 3DSS-VLG, a weakly supervised approach for 3D Semantic Segmentation with 2D Vision-Language Guidance, an alternative approach that a 3D model predicts den…
GradBias: Unveiling Word Influence on Bias in Text-to-Image Generative Models
Moreno D'Incà, Elia Peruzzo, Massimiliano Mancini +3
Recent progress in Text-to-Image (T2I) generative models has enabled high-quality image generation. As performance and accessibility increase, these models are gaining significant…
Consistency-Aware Anchor Pyramid Network for Crowd Localization
Xinyan Liu, Guorong Li, Yuankai Qi +4
Crowd localization aims to predict the spatial position of humans in a crowd scenario. We observe that the performance of existing methods is challenged from two aspects: (i) ranki…