collaborators

5 papers

cs.CV2025

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Haoji Zhang, Yiqin Wang, Yansong Tang +3

Benefiting from the advances in large language models and cross-modal alignment, existing multimodal large language models have achieved prominent performance in image and short vi…

cs.CV2025

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Zhongwei Ren, Yunchao Wei, Xun Guo +4

This work explores whether a deep generative model can learn complex knowledge solely from visual input, in contrast to the prevalent focus on text-based models like large language…

cs.CV2025

Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens

Fan Ma, Xiaojie Jin, Heng Wang +3

Recent advances in large video-language models have displayed promising outcomes in video comprehension. Current approaches straightforwardly convert video into language tokens and…

cs.CV2024

Hierarchical Memory for Long Video QA

Yiqin Wang, Haoji Zhang, Yansong Tang +4

This paper describes our champion solution to the LOVEU Challenge @ CVPR'24, Track 1 (Long Video VQA). Processing long sequences of visual tokens is computationally expensive and m…

cs.CV2024

Noise Self-Regression: A New Learning Paradigm to Enhance Low-Light Images Without Task-Related Data

Zhao Zhang, Suiyi Zhao, Xiaojie Jin +4

Deep learning-based low-light image enhancement (LLIE) is a task of leveraging deep neural networks to enhance the image illumination while keeping the image content unchanged. Fro…