activity
20242026
most citedStreaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

1 citations · 2 across the 20 of their papers we have counts for

collaborators
Showing 2025 · cs.CVShow all

6 papers · 2 filters

cs.CV2025

From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction

Zhida Zhao, Talas Fu, Yifan Wang +2

Despite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and d…

cs.CV2025

Learning Universal Features for Generalizable Image Forgery Localization

Hengrun Zhao, Yunzhi Zhuge, Yifan Wang +3

In recent years, advanced image editing and generation methods have rapidly evolved, making detecting and locating forged image content increasingly challenging. Most existing imag…

cs.CV2025

Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion

Songsong Yu, Yuxin Chen, Zhongang Qi +5

With the rapid proliferation of 3D devices and the shortage of 3D content, stereo conversion is attracting increasing attention. Recent works introduce pretrained Diffusion Models…

cs.CV2025★ 1 cited

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

Haomiao Xiong, Zongxin Yang, Jiazuo Yu +4

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, curre…

cs.CV2025

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding

Haomiao Xiong, Yunzhi Zhuge, Jiawen Zhu +2

Multi-modal Large Language Models (MLLMs) exhibit impressive capabilities in 2D tasks, yet encounter challenges in discerning the spatial positions, interrelations, and causal logi…

cs.CV2025

AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation

Sitong Gong, Yunzhi Zhuge, Lu Zhang +4

The essence of audio-visual segmentation (AVS) lies in locating and delineating sound-emitting objects within a video stream. While Transformer-based methods have shown promise, th…