3 papers
cs.CV2025
Diffusion Is Your Friend in Show, Suggest and Tell
Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi
Diffusion Denoising models demonstrated impressive results across generative Computer Vision tasks, but they still fail to outperform standard autoregressive solutions in the discr…
cs.CV2024
Shifted Window Fourier Transform And Retention For Image Captioning
Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi
Image Captioning is an important Language and Vision task that finds application in a variety of contexts, ranging from healthcare to autonomous vehicles. As many real-world applic…
cs.CL2024
Bidirectional Awareness Induction in Autoregressive Seq2Seq Models
Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi
Autoregressive Sequence-To-Sequence models are the foundation of many Deep Learning achievements in major research fields such as Vision and Natural Language Processing. Despite th…