most citedLong Context Transfer from Language to Vision

2 citations · 2 across the 1 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey

Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang +10

Detecting out-of-distribution (OOD) samples is crucial for ensuring the safety of machine learning systems and has shaped the field of OOD detection. Meanwhile, several other probl…

cs.CV20242 cited

Long Context Transfer from Language to Vision

Peiyuan Zhang, Kaichen Zhang, Bo Li +7

Video sequences offer valuable temporal information, but existing large multimodal models (LMMs) fall short in understanding extremely long videos. Many works address this by reduc…

cs.CV2024

4D Panoptic Scene Graph Generation

Jingkang Yang, Jun Cen, Wenxuan Peng +6

We are living in a three-dimensional space while moving forward through a fourth dimension: time. To allow artificial intelligence to develop a comprehensive understanding of such…

cs.CV2024

Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models

Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang +7

This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed . Multiple…

cs.CV2024

Towards Language-Driven Video Inpainting via Multimodal Large Language Models

Jianzong Wu, Xiangtai Li, Chenyang Si +8

We introduce a new task -- language-driven video inpainting, which uses natural language instructions to guide the inpainting process. This approach overcomes the limitations of tr…