most citedUnveiling Cross Modality Bias in Visual Question Answering: A Causal View with Possible Worlds VQA

2 citations · 2 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2024

OSCaR: Object State Captioning and State Change Representation

Nguyen Nguyen, Jing Bi, Ali Vosoughi +3

The capability of intelligent models to extrapolate and comprehend changes in object states is a crucial yet demanding aspect of AI research, particularly through the lens of human…

cs.MM2024

Learning Audio Concepts from Counterfactual Natural Language

Ali Vosoughi, Luca Bondi, Ho-Hsiang Wu +1

Conventional audio classification relied on predefined classes, lacking the ability to learn from free-form text. Recent methods unlock learning joint audio-text embeddings from ra…

cs.CV2023

Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation

Yiyang Su, Ali Vosoughi, Shijian Deng +2

The audio-visual sound separation field assumes visible sources in videos, but this excludes invisible sounds beyond the camera's view. Current methods struggle with such sounds la…

cs.CL2023

MISAR: A Multimodal Instructional System with Augmented Reality

Jing Bi, Nguyen Manh Nguyen, Ali Vosoughi +1

Augmented reality (AR) requires the seamless integration of visual, auditory, and linguistic channels for optimized human-computer interaction. While auditory and visual inputs fac…

cs.CV20232 cited

Unveiling Cross Modality Bias in Visual Question Answering: A Causal View with Possible Worlds VQA

Ali Vosoughi, Shijian Deng, Songyang Zhang +3

To increase the generalization capability of VQA systems, many recent studies have tried to de-bias spurious language or vision associations that shortcut the question or image to…

cs.CV2023

Cross Modal Global Local Representation Learning from Radiology Reports and X-Ray Chest Images

Nathan Hadjiyski, Ali Vosoughi, Axel Wismueller

Deep learning models can be applied successfully in real-work problems; however, training most of these models requires massive data. Recent methods use language and vision, but un…