33 papers
XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition
Kun-Yu Lin, Henghui Ding, Jia-Run Du +6
Inspired by the impressive success of image-text foundation models, recent works have proposed to adapt these foundation models to video data, leading to efficient and effective vi…
AnyBokeh: Physics-Guided Any-to-Any Bokeh Editing with Optical Fingerprint Transfer
Xinyu Hou, Xiaoming Li, Zongsheng Yue +1
Depth-of-field control is a fundamental tool in photography, yet post-capture bokeh editing from a single image remains challenging. A practical editor should handle images capture…
Generate Your Talking Avatar from Video Reference
Zujin Guo, Zhenhui Ye, Yi Ren +4
Existing talking avatar methods typically adopt an image-to-video pipeline conditioned on a static reference image within the same scene as the target generation. This restricted,…
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
Kang Liao, Size Wu, Zhonghua Wu +6
Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric m…
Point-In-Context: Understanding Point Cloud via In-Context Learning
Mengyuan Liu, Zhongbin Fang, Xia Li +4
The rise of large-scale models has catalyzed in-context learning as a powerful approach for multitasking, particularly in natural language and image processing. However, its applic…
Next Visual Granularity Generation
Yikai Wang, Zhouxia Wang, Zhonghua Wu +3
We propose a novel approach to image generation by decomposing an image into a structured sequence, where each element in the sequence shares the same spatial resolution but differ…