2 citations · 4 across the 25 of their papers we have counts for
27 papers
WorldSculpt: Generating Compositional Worlds from Grounded Videos
Muyao Niu, Jixuan He, Ruihan Yu +9
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of indi…
SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language Models
Caoyuan Ma, Tian Gu, Wenpu Liu +11
Existing safety alignment methods for vision-language models usually modify the model behavior globally: once the safety parameters are trained or loaded, they participate in both…
Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation
Guixu Lin, Yuyang Yu, Xiang Ji +6
Latent diffusion models have recently advanced video frame interpolation by synthesizing intermediate frames between input images. However, handling large temporal gaps and complex…
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
Caoyuan Ma, Wenpu Liu, Weichu Xie +12
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propos…
Optical Flow from Photons
Wendi Liu, Weichao Zeng, Weihang Ran +2
Optical flow remains challenging in high-speed and low-light scenes, where the limited frame rate and sensitivity of conventional cameras lead to motion blur and underexposure. Sin…
Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark
Muyao Niu, Mingze Ma, Yifan Zhan +5
Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NIR) with noisy RGB to achieve improved enhancement, yet most met…