activity
20242026
collaborators

6 papers

cs.CV2026

From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification

Ke Zhang, Xiangchen Zhao, Yunjie Tian +3

Conventional video classification models, acting as effective imitators, excel in scenarios with homogeneous data distributions. However, real-world applications often present an o…

cs.CV2026

Few-Step Diffusion Sampling Through Instance-Aware Discretizations

Liangyu Yuan, Ruoyu Wang, Tong Zhao +4

Diffusion and flow matching models generate high-fidelity data by simulating paths defined by Ordinary or Stochastic Differential Equations (ODEs/SDEs), starting from a tractable p…

cs.CV2026

Shot-Aware Frame Sampling for Video Understanding

Mengyu Zhao, Di Fu, Yongyu Xie +4

Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when…

cs.CV2026

CleanStyle: Plug-and-Play Style Conditioning Purification for Text-to-Image Stylization

Xiaoman Feng, Mingkun Lei, Yang Wang +2

Style transfer in diffusion models enables controllable visual generation by injecting the style of a reference image. However, recent encoder-based methods, while efficient and tu…

cs.CV2025

RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification

Meilong Xu, Di Fu, Jiaxing Zhang +7

Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particul…

cs.CV2024

Shot Segmentation Based on Von Neumann Entropy for Key Frame Extraction

Xueqing Zhang, Di Fu, Naihao Liu

Video key frame extraction is important in various fields, such as video summary, retrieval, and compression. Therefore, we suggest a video key frame extraction algorithm based on…