3 citations · 4 across the 4 of their papers we have counts for
3 papers · 1 filter
CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning
SeongJun Jeong, Minjoon Jung, Woo Suk Choi +2
Vision-language models (VLMs) demonstrate strong performance across compositional reasoning benchmarks, which require reasoning over semantic perturbations of objects, attributes,…
Instruction-tuned Self-Questioning Framework for Multimodal Reasoning
You-Won Jang, Yu-Jung Heo, Jaeseok Kim +3
The field of vision-language understanding has been actively researched in recent years, thanks to the development of Large Language Models~(LLMs). However, it still needs help wit…
Mounting Video Metadata on Transformer-based Language Model for Open-ended Video Question Answering
Donggeon Lee, Seongho Choi, Youwon Jang +1
Video question answering has recently received a lot of attention from multimodal video researchers. Most video question answering datasets are usually in the form of multiple-choi…