7 papers
Revisiting Salient Object Detection from an Observer-Centric Perspective
Fuxi Zhang, Yifan Wang, Hengrun Zhao +7
Salient object detection is inherently a subjective problem, as observers with different priors may perceive different objects as salient. However, existing methods predominantly f…
AR-MOT: Autoregressive Multi-object Tracking
Lianjie Jia, Yuhan Wu, Binghao Ran +3
As multi-object tracking (MOT) tasks continue to evolve toward more general and multi-modal scenarios, the rigid and task-specific architectures of existing MOT methods increasingl…
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
Zhida Zhao, Talas Fu, Yifan Wang +2
Despite remarkable progress in driving world models, their potential for autonomous systems remains largely untapped: the world models are mostly learned for world simulation and d…
Learning Universal Features for Generalizable Image Forgery Localization
Hengrun Zhao, Yunzhi Zhuge, Yifan Wang +3
In recent years, advanced image editing and generation methods have rapidly evolved, making detecting and locating forged image content increasingly challenging. Most existing imag…
Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion
Songsong Yu, Yuxin Chen, Zhongang Qi +5
With the rapid proliferation of 3D devices and the shortage of 3D content, stereo conversion is attracting increasing attention. Recent works introduce pretrained Diffusion Models…
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
Sitong Gong, Yunzhi Zhuge, Lu Zhang +4
The essence of audio-visual segmentation (AVS) lies in locating and delineating sound-emitting objects within a video stream. While Transformer-based methods have shown promise, th…