6 papers
From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification
Ke Zhang, Xiangchen Zhao, Yunjie Tian +3
Conventional video classification models, acting as effective imitators, excel in scenarios with homogeneous data distributions. However, real-world applications often present an o…
Few-Step Diffusion Sampling Through Instance-Aware Discretizations
Liangyu Yuan, Ruoyu Wang, Tong Zhao +4
Diffusion and flow matching models generate high-fidelity data by simulating paths defined by Ordinary or Stochastic Differential Equations (ODEs/SDEs), starting from a tractable p…
Shot-Aware Frame Sampling for Video Understanding
Mengyu Zhao, Di Fu, Yongyu Xie +4
Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when…
CleanStyle: Plug-and-Play Style Conditioning Purification for Text-to-Image Stylization
Xiaoman Feng, Mingkun Lei, Yang Wang +2
Style transfer in diffusion models enables controllable visual generation by injecting the style of a reference image. However, recent encoder-based methods, while efficient and tu…
RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
Meilong Xu, Di Fu, Jiaxing Zhang +7
Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particul…
Shot Segmentation Based on Von Neumann Entropy for Key Frame Extraction
Xueqing Zhang, Di Fu, Naihao Liu
Video key frame extraction is important in various fields, such as video summary, retrieval, and compression. Therefore, we suggest a video key frame extraction algorithm based on…