6 papers
Adaptive Keyframe Sampling for Long Video Understanding
Xi Tang, Jihao Qiu, Lingxi Xie +3
Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. Howev…
Rethinking Sampling Strategies for Unsupervised Person Re-identification
Xumeng Han, Xuehui Yu, Guorong Li +5
Unsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that samp…
Depth-guided Texture Diffusion for Image Semantic Segmentation
Wei Sun, Yuan Li, Qixiang Ye +2
Depth information provides valuable insights into the 3D structure especially the outline of objects, which can be utilized to improve the semantic segmentation tasks. However, a n…
Correspondence-Guided SfM-Free 3D Gaussian Splatting for NVS
Wei Sun, Xiaosong Zhang, Fang Wan +4
Novel View Synthesis (NVS) without Structure-from-Motion (SfM) pre-processed camera poses--referred to as SfM-free methods--is crucial for promoting rapid response capabilities and…
Self-supervised Feature-Gate Coupling for Dynamic Network Pruning
Mengnan Shi, Chang Liu, Jianbin Jiao +1
Gating modules have been widely explored in dynamic network pruning to reduce the run-time computational cost of deep neural networks while preserving the representation of feature…
Uncertainty-guided Optimal Transport in Depth Supervised Sparse-View 3D Gaussian
Wei Sun, Qi Zhang, Yanzhao Zhou +3
3D Gaussian splatting has demonstrated impressive performance in real-time novel view synthesis. However, achieving successful reconstruction from RGB images generally requires mul…