30 citations · 57 across the 8 of their papers we have counts for
3 papers · 1 filter
Diffusion Is Your Friend in Show, Suggest and Tell
Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi
Diffusion Denoising models demonstrated impressive results across generative Computer Vision tasks, but they still fail to outperform standard autoregressive solutions in the discr…
Shifted Window Fourier Transform And Retention For Image Captioning
Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi
Image Captioning is an important Language and Vision task that finds application in a variety of contexts, ranging from healthcare to autonomous vehicles. As many real-world applic…
A request for clarity over the End of Sequence token in the Self-Critical Sequence Training
Jia Cheng Hu, Roberto Cavicchioli, Alessandro Capotondi
The Image Captioning research field is currently compromised by the lack of transparency and awareness over the End-of-Sequence token (<Eos>) in the Self-Critical Sequence Training…