activity
20242026
collaborators

6 papers

cs.CV2026

Advancing Open-source World Models

Robbyant Team, Zelin Gao, Qiuyu Wang +21

We present LingBot-World, an open-sourced world simulator stemming from video generation. Positioned as a top-tier world model, LingBot-World offers the following features. (1) It…

cs.CV2025

Aligned Better, Listen Better for Audio-Visual Large Language Models

Yuxin Guo, Shuailei Ma, Shijie Ma +7

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large la…

cs.CV2025

Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning

Fan Lu, Wei Wu, Kecheng Zheng +7

Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have dev…

cs.CV2025

Learning Visual Generative Priors without Text

Shuailei Ma, Kecheng Zheng, Ying Wei +7

Although text-to-image (T2I) models have recently thrived as visual generative priors, their reliance on high-quality text-image pairs makes scaling up expensive. We argue that gra…

cs.CV2024

LoTLIP: Improving Language-Image Pre-training for Long Text Understanding

Wei Wu, Kecheng Zheng, Shuailei Ma +7

Understanding long text is of great demands in practice but beyond the reach of most language-image pre-training (LIP) models. In this work, we empirically confirm that the key rea…

cs.CV2024

CoReS: Orchestrating the Dance of Reasoning and Segmentation

Xiaoyi Bao, Siyang Sun, Shuailei Ma +5

The reasoning segmentation task, which demands a nuanced comprehension of intricate queries to accurately pinpoint object regions, is attracting increasing attention. However, Mult…