activity
20222025
most citedQwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

149 citations · 385 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CL20252 cited

Qwen3-Omni Technical Report

Jin Xu, Zhifang Guo, Hangrui Hu +35

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relat…

cs.CL202512 cited

Qwen2.5-Omni Technical Report

Jin Xu, Zhifang Guo, Jinzheng He +11

In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously gene…

cs.CL202460 cited

Qwen2 Technical Report

An Yang, Baosong Yang, Binyuan Hui +59

This report introduces the Qwen2 series, the latest addition to our large language models and large multimodal models. We release a comprehensive suite of foundational and instruct…

cs.CV2024

GD^2-NeRF: Generative Detail Compensation via GAN and Diffusion for One-shot Generalizable Neural Radiance Fields

Xiao Pan, Zongxin Yang, Shuai Bai +1

In this paper, we focus on the One-shot Novel View Synthesis (O-NVS) task which targets synthesizing photo-realistic novel views given only one reference image per scene. Previous…

cs.CL2023110 cited

Qwen Technical Report

Jinze Bai, Shuai Bai, Yunfei Chu +45

Large language models (LLMs) have revolutionized the field of artificial intelligence, enabling natural language processing tasks that were previously thought to be exclusive to hu…

cs.CV20237 cited

TouchStone: Evaluating Vision-Language Models by Language Models

Shuai Bai, Shusheng Yang, Jinze Bai +6

Large vision-language models (LVLMs) have recently witnessed rapid advancements, exhibiting a remarkable capacity for perceiving, understanding, and processing visual information b…