most citedQwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

149 citations · 194 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV2024

Learning Visual Grounding from Generative Vision and Language Model

Shijie Wang, Dahun Kim, Ali Taalimi +2

Visual grounding tasks aim to localize image regions based on natural language references. In this work, we explore whether generative VLMs predominantly trained on image-text data…

cs.CV2024

Adaptive Fusion of Single-View and Multi-View Depth for Autonomous Driving

JunDa Cheng, Wei Yin, Kaixuan Wang +3

Multi-view depth estimation has achieved impressive performance over various benchmarks. However, almost all current multi-view systems rely on given ideal camera poses, which are…

cs.CV2023

Object-centric Video Representation for Long-term Action Anticipation

Ce Zhang, Changcheng Fu, Shijie Wang +4

This paper focuses on building object-centric representations for long-term action anticipation in videos. Our key motivation is that objects provide important cues to recognize an…

cs.CV2023149 cited

Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Jinze Bai, Shuai Bai, Shusheng Yang +6

In this work, we introduce the Qwen-VL series, a set of large-scale vision-language models (LVLMs) designed to perceive and understand both texts and images. Starting from the Qwen…

cs.AI20231 cited

A Novel Multi-Agent Deep RL Approach for Traffic Signal Control

Shijie Wang, Shangbo Wang

As travel demand increases and urban traffic condition becomes more complicated, applying multi-agent deep reinforcement learning (MARL) to traffic signal control becomes one of th…

cs.CV202342 cited

ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

Peng Wang, Shijie Wang, Junyang Lin +5

In this work, we explore a scalable way for building a general representation model toward unlimited modalities. We release ONE-PEACE, a highly extensible model with 4B parameters…