3 papers
cs.CV2024
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
João Bordalo, Vasco Ramos, Rodrigo Valério +5
Multistep instructions, such as recipes and how-to guides, greatly benefit from visual aids, such as a series of images that accompany the instruction steps. While Large Language M…
cs.CL2023
TWIZ-v2: The Wizard of Multimodal Conversational-Stimulus
Rafael Ferreira, Diogo Tavares, Diogo Silva +6
In this report, we describe the vision, challenges, and scientific contributions of the Task Wizard team, TWIZ, in the Alexa Prize TaskBot Challenge 2022. Our vision, is to build T…
cs.CV2023
Transferring Visual Attributes from Natural Language to Verified Image Generation
Rodrigo Valerio, Joao Bordalo, Michal Yarom +3
Text to image generation methods (T2I) are widely popular in generating art and other creative artifacts. While visual hallucinations can be a positive factor in scenarios where cr…