1 citations · 1 across the 1 of their papers we have counts for
4 papers
Show and Guide: Instructional-Plan Grounded Vision and Language Model
Diogo Glória-Silva, David Semedo, João Magalhães
Guiding users through complex procedural plans is an inherently multimodal task in which having visually illustrated plan steps is crucial to deliver an effective plan guidance. Ho…
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
João Bordalo, Vasco Ramos, Rodrigo Valério +5
Multistep instructions, such as recipes and how-to guides, greatly benefit from visual aids, such as a series of images that accompany the instruction steps. While Large Language M…
Plan-Grounded Large Language Models for Dual Goal Conversational Settings
Diogo Glória-Silva, Rafael Ferreira, Diogo Tavares +2
Training Large Language Models (LLMs) to follow user instructions has been shown to supply the LLM with ample capacity to converse fluently while being aligned with humans. Yet, it…
TWIZ-v2: The Wizard of Multimodal Conversational-Stimulus
Rafael Ferreira, Diogo Tavares, Diogo Silva +6
In this report, we describe the vision, challenges, and scientific contributions of the Task Wizard team, TWIZ, in the Alexa Prize TaskBot Challenge 2022. Our vision, is to build T…