activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs

Baiyang Song, Jun Peng, Yuxin Zhang +3

Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a video as a sequence of static fra…

cs.CV2025

Improving Video Generation with Human Feedback

Jie Liu, Gongye Liu, Jiajun Liang +14

Video generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this w…

cs.CV2025

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Jiahao Hu, Tianxiong Zhong, Xuebo Wang +5

Diffusion-based image editing models have made remarkable progress in recent years. However, achieving high-quality video editing remains a significant challenge. One major hurdle…

cs.CV2025

Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Qiuheng Wang, Yukai Shi, Jiarong Ou +10

With the continuous progress of visual generation technologies, the scale of video datasets has grown exponentially. The quality of these datasets plays a pivotal role in the perfo…

cs.CV2024

Focus On What Matters: Separated Models For Visual-Based RL Generalization

Di Zhang, Bowen Lv, Hai Zhang +7

A primary challenge for visual-based Reinforcement Learning (RL) is to generalize effectively across unseen environments. Although previous studies have explored different auxiliar…