most citedA Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation

1 citations · 1 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV2025

SegMo: Segment-aligned Text to 3D Human Motion Generation

Bowen Dang, Lin Wu, Xiaohang Yang +2

Generating 3D human motions from textual descriptions is an important research problem with broad applications in video games, virtual reality, and augmented reality. Recent method…

cs.CV2025

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models

Mingyu Fu, Wei Suo, Ji Ma +3

Despite the great success of Large Vision Language Models (LVLMs), their high computational cost severely limits their broad applications. The computational cost of LVLMs mainly st…

cs.CV20251 cited

A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation

Qing Zhong, Peng-Tao Jiang, Wen Wang +3

Contemporary Video Instance Segmentation (VIS) methods typically adhere to a pre-train then fine-tune regime, where a segmentation model trained on images is fine-tuned on videos.…

cs.CV2025

Unlocking Generalization Power in LiDAR Point Cloud Registration

Zhenxuan Zeng, Qiao Wu, Xiyu Zhang +5

In real-world environments, a LiDAR point cloud registration method with robust generalization capabilities (across varying distances and datasets) is crucial for ensuring safety i…

cs.CV2025

Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding

Wei Suo, Lijun Zhang, Mengyang Sun +3

Large Vision-Language Models (LVLMs) have obtained impressive performance in visual content understanding and multi-modal reasoning. Unfortunately, these large models suffer from s…

cs.CV2024

A Deep Semantic Segmentation Network with Semantic and Contextual Refinements

Zhiyan Wang, Deyin Liu, Lin Yuanbo Wu +3

Semantic segmentation is a fundamental task in multimedia processing, which can be used for analyzing, understanding, editing contents of images and videos, among others. To accele…