7 citations · 7 across the 3 of their papers we have counts for
3 papers
VISREAS: Complex Visual Reasoning with Unanswerable Questions
Syeda Nahida Akter, Sangwu Lee, Yingshan Chang +2
Verifying a question's validity before answering is crucial in real-world applications, where users may provide imperfect instructions. In this scenario, an ideal model should addr…
Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination
Syeda Nahida Akter, Aman Madaan, Sangwu Lee +2
The potential of Vision-Language Models (VLMs) often remains underutilized in handling complex text-based problems, particularly when these problems could benefit from visual repre…
TextMI: Textualize Multimodal Information for Integrating Non-verbal Cues in Pre-trained Language Models
Md Kamrul Hasan, Md Saiful Islam, Sangwu Lee +4
Pre-trained large language models have recently achieved ground-breaking performance in a wide variety of language understanding tasks. However, the same model can not be applied t…