4 papers
Reducing Bias and Variance: Generative Semantic Guidance and Bi-Layer Ensemble for Image Clustering
Feijiang Li, Zhenxiong Li, Jieting Wang +3
Image clustering aims to partition unlabeled image datasets into distinct groups. A core aspect of this task is constructing and leveraging prior knowledge to guide the clustering…
OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models
Yuchen Deng, Zidang Cai, Hai-Tao Zheng +3
Omnimodal large language models (Omni-LLMs) show strong capability in audio-video understanding, but their practical deployment remains limited by high inference cost of long video…
Beyond Boundary Frames: Talking-Head Inbetweening via Context-Aware Motion Modeling
Yuchen Deng, Xiuyang Wu, Hai-Tao Zheng +4
Existing talking-head generation methods primarily target open-ended generation rather than bridging two existing video segments. In this paper, we study talking-head inbetweening,…
Flow Intelligence: Robust Feature Matching via Temporal Signature Correlation
Jie Wang, Chen Ye Gan, Caoqi Wei +2
Feature matching across video streams remains a cornerstone challenge in computer vision. Increasingly, robust multimodal matching has garnered interest in robotics, surveillance,…