28 papers
OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films
Xin Lu, Zihao Fan, Mingchen Zhong +3
Historical films suffer from co-occurring visual and audio degradations---blur, noise, flicker, hiss, clipping, and dropout---yet existing methods restore each modality independent…
Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models
Yawen Shao, Jie Xiao, Kai Zhu +6
Reinforcement learning (RL) holds immense promise for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, progress is fundamentally constraine…
Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset
Senyan Xu, Zhijing Sun, Kean Liu +5
Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking the essential global illumina…
EventGait: Towards Robust Gait Recognition with Event Streams
Senyan Xu, Shuai Chen, Chuanfu Shen +4
Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensitivity of conventional camera…
IR-Flow: Bridging Discriminative and Generative Image Restoration via Rectified Flow
Zihao Fan, Xin Lu, Jie Xiao +3
In image restoration, single-step discriminative mappings often lack fine details via expectation learning, whereas generative paradigms suffer from inefficient multi-step sampling…
GS-STVSR: Ultra-Efficient Continuous Spatio-Temporal Video Super-Resolution via 2D Gaussian Splatting
Mingyu Shi, Xin Di, Long Peng +8
Continuous Spatio-Temporal Video Super-Resolution (C-STVSR) aims to simultaneously enhance the spatial resolution and frame rate of videos by arbitrary scale factors, offering grea…