collaborators

6 papers

cs.CV2026

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

Zijun Wang, Panwen Hu, Jing Wang +7

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent resul…

cs.LG2025

A Modular, Data-Free Pipeline for Multi-Label Intention Recognition in Transportation Agentic AI Applications

Xiaocai Zhang, Hur Lim, Ke Wang +5

In this study, a modular, data-free pipeline for multi-label intention recognition is proposed for agentic AI applications in transportation. Unlike traditional intent recognition…

cs.CV2025

IF-VidCap: Can Video Caption Models Follow Instructions?

Shihao Li, Yuanxing Zhang, Jiangtao Wu +20

Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions…

cs.CV2025

Kwai Keye-VL 1.5 Technical Report

Biao Yang, Bin Wen, Boyang Ding +58

In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Mode…

cs.CV2025

Kwai Keye-VL Technical Report

Kwai Keye Team, Biao Yang, Bin Wen +57

While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form vi…

cs.CV2025

RepText: Rendering Visual Text via Replicating

Haofan Wang, Yujia Xu, Yimeng Li +5

Although contemporary text-to-image generation models have achieved remarkable breakthroughs in producing visually appealing images, their capacity to generate precise and flexible…