collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2025

SegMo: Segment-aligned Text to 3D Human Motion Generation

Bowen Dang, Lin Wu, Xiaohang Yang +2

Generating 3D human motions from textual descriptions is an important research problem with broad applications in video games, virtual reality, and augmented reality. Recent method…

cs.CV2025

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models

Mingyu Fu, Wei Suo, Ji Ma +3

Despite the great success of Large Vision Language Models (LVLMs), their high computational cost severely limits their broad applications. The computational cost of LVLMs mainly st…

cs.CV2025

Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models

Wei Suo, Ji Ma, Mengyang Sun +3

Although Large Vision-Language Models (LVLMs) have achieved impressive results, their high computational costs pose a significant barrier to wide application. To enhance inference…

cs.CV2025

A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation

Qing Zhong, Peng-Tao Jiang, Wen Wang +3

Contemporary Video Instance Segmentation (VIS) methods typically adhere to a pre-train then fine-tune regime, where a segmentation model trained on images is fine-tuned on videos.…

cs.CV2025

Unlocking Generalization Power in LiDAR Point Cloud Registration

Zhenxuan Zeng, Qiao Wu, Xiyu Zhang +5

In real-world environments, a LiDAR point cloud registration method with robust generalization capabilities (across varying distances and datasets) is crucial for ensuring safety i…

cs.CV2025

Octopus: Alleviating Hallucination via Dynamic Contrastive Decoding

Wei Suo, Lijun Zhang, Mengyang Sun +3

Large Vision-Language Models (LVLMs) have obtained impressive performance in visual content understanding and multi-modal reasoning. Unfortunately, these large models suffer from s…