collaborators

6 papers

cs.CV2025

Adaptive Keyframe Sampling for Long Video Understanding

Xi Tang, Jihao Qiu, Lingxi Xie +3

Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. Howev…

cs.CV2024

Rethinking Sampling Strategies for Unsupervised Person Re-identification

Xumeng Han, Xuehui Yu, Guorong Li +5

Unsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that samp…

cs.CV2024

Depth-guided Texture Diffusion for Image Semantic Segmentation

Wei Sun, Yuan Li, Qixiang Ye +2

Depth information provides valuable insights into the 3D structure especially the outline of objects, which can be utilized to improve the semantic segmentation tasks. However, a n…

cs.CV2024

Correspondence-Guided SfM-Free 3D Gaussian Splatting for NVS

Wei Sun, Xiaosong Zhang, Fang Wan +4

Novel View Synthesis (NVS) without Structure-from-Motion (SfM) pre-processed camera poses--referred to as SfM-free methods--is crucial for promoting rapid response capabilities and…

cs.CV2024

Self-supervised Feature-Gate Coupling for Dynamic Network Pruning

Mengnan Shi, Chang Liu, Jianbin Jiao +1

Gating modules have been widely explored in dynamic network pruning to reduce the run-time computational cost of deep neural networks while preserving the representation of feature…

cs.CV2024

Uncertainty-guided Optimal Transport in Depth Supervised Sparse-View 3D Gaussian

Wei Sun, Qi Zhang, Yanzhao Zhou +3

3D Gaussian splatting has demonstrated impressive performance in real-time novel view synthesis. However, achieving successful reconstruction from RGB images generally requires mul…