17 papers
A Bayesian Proof of the Bernoulli Theorem
Jingbo Liu, Ilias Zadik
We give a new proof of the Bernoulli theorem, conjectured by Talagrand and proved in the seminal work of Bednorz and Latała. Our approach is based on information-theoretic ideas: l…
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
Caoyuan Ma, Wenpu Liu, Weichu Xie +12
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propos…
Optical Flow from Photons
Wendi Liu, Weichao Zeng, Weihang Ran +2
Optical flow remains challenging in high-speed and low-light scenes, where the limited frame rate and sensitivity of conventional cameras lead to motion blur and underexposure. Sin…
Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark
Muyao Niu, Mingze Ma, Yifan Zhan +5
Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NIR) with noisy RGB to achieve improved enhancement, yet most met…
Surprise Forcing: What to Remember, When to Skip in Long Video Generation
Shuwei Shi, Zhen Li, Muyao Niu +4
Streaming autoregressive diffusion makes minute-scale video synthesis practical, but its bounded context and fixed denoising schedule allocate resources uniformly across a highly n…
MoECodec: Image Compression for joint human and machine perception via Mixture-of-Experts
Jiancheng Zhao, Xiang Ji, Yifan Zhan +2
Image compression for machines calls for a unified codec that serves multiple downstream vision tasks. Existing approaches either adopt task-specific end-to-end designs, raising pa…