2 citations · 2 across the 1 of their papers we have counts for
5 papers · 1 filter
Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang +10
Detecting out-of-distribution (OOD) samples is crucial for ensuring the safety of machine learning systems and has shaped the field of OOD detection. Meanwhile, several other probl…
Long Context Transfer from Language to Vision
Peiyuan Zhang, Kaichen Zhang, Bo Li +7
Video sequences offer valuable temporal information, but existing large multimodal models (LMMs) fall short in understanding extremely long videos. Many works address this by reduc…
4D Panoptic Scene Graph Generation
Jingkang Yang, Jun Cen, Wenxuan Peng +6
We are living in a three-dimensional space while moving forward through a fourth dimension: time. To allow artificial intelligence to develop a comprehensive understanding of such…
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang +7
This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed . Multiple…
Towards Language-Driven Video Inpainting via Multimodal Large Language Models
Jianzong Wu, Xiangtai Li, Chenyang Si +8
We introduce a new task -- language-driven video inpainting, which uses natural language instructions to guide the inpainting process. This approach overcomes the limitations of tr…