2 papers
cs.CV2020
VICTR: Visual Information Captured Text Representation for Text-to-Image Multimodal Tasks
Soyeon Caren Han, Siqu Long, Siwen Luo +2
Text-to-image multimodal tasks, generating/retrieving an image from a given text description, are extremely challenging tasks since raw text descriptions cover quite limited inform…
cs.CV2020
REXUP: I REason, I EXtract, I UPdate with Structured Compositional Reasoning for Visual Question Answering
Siwen Luo, Soyeon Caren Han, Kaiyuan Sun +1
Visual question answering (VQA) is a challenging multi-modal task that requires not only the semantic understanding of both images and questions, but also the sound perception of a…