activity
20212024
most citedThe Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

172 citations · 305 across the 12 of their papers we have counts for

collaborators

12 papers

cs.CV2024

Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition

Jielin Qiu, William Han, Winfred Wang +6

Open-domain real-world entity recognition is essential yet challenging, involving identifying various entities in diverse environments. The lack of a suitable evaluation dataset ha…

cs.CV202310 cited

GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

An Yan, Zhengyuan Yang, Wanrong Zhu +9

We present MM-Navigator, a GPT-4V-based agent for the smartphone graphical user interface (GUI) navigation task. MM-Navigator can interact with a smartphone screen as human users,…

cs.CV20231 cited

OpenLEAF: Open-Domain Interleaved Image-Text Generation and Evaluation

Jie An, Zhengyuan Yang, Linjie Li +5

This work investigates a challenging task named open-domain interleaved image-text generation, which generates interleaved texts and images following an input query. We propose a n…

cs.CV202310 cited

MM-VID: Advancing Video Understanding with GPT-4V(ision)

Kevin Lin, Faisal Ahmed, Linjie Li +9

We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video unders…

cs.CV20231 cited

DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual Design

Kevin Lin, Zhengyuan Yang, Linjie Li +2

We introduce DEsignBench, a text-to-image (T2I) generation benchmark tailored for visual design scenarios. Recent T2I models like DALL-E 3 and others, have demonstrated remarkable…

cs.CV2023172 cited

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Zhengyuan Yang, Linjie Li, Kevin Lin +4

Large multimodal models (LMMs) extend large language models (LLMs) with multi-sensory skills, such as visual understanding, to achieve stronger generic intelligence. In this paper,…