activity
20202025
most citedL2G: A Simple Local-to-Global Knowledge Transfer Framework for Weakly Supervised Semantic Segmentation

10 citations · 29 across the 29 of their papers we have counts for

collaborators
Showing cs.CVShow all

32 papers · 1 filter

cs.CV2025

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models

Quan-Sheng Zeng, Yunheng Li, Qilong Wang +4

Visual token compression is critical for Large Vision-Language Models (LVLMs) to efficiently process high-resolution inputs. Existing methods that typically adopt fixed compression…

cs.CV2025

Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment

Zhuoxuan Cai, Jian Zhang, Xinbin Yuan +7

Recent studies demonstrate that multimodal large language models (MLLMs) can proficiently evaluate visual quality through interpretable assessments. However, existing approaches ty…

cs.CV2025

HyperMotionX: The Dataset and Benchmark with DiT-Based Pose-Guided Human Image Animation of Complex Motions

Shuolin Xu, Siming Zheng, Ziyi Wang +7

Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods…

cs.CV2025

MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on

Guangyuan Li, Siming Zheng, Hao Zhang +6

Video Virtual Try-On (VVT) aims to synthesize garments that appear natural across consecutive video frames, capturing both their dynamics and interactions with human motion. Despit…

cs.CV2025

Photography Perspective Composition: Towards Aesthetic Perspective Recommendation

Lujian Yao, Siming Zheng, Xinbin Yuan +5

Traditional photography composition approaches are dominated by 2D cropping-based methods. However, these methods fall short when scenes contain poorly arranged subjects. Professio…

cs.CV2025★ 1 cited

A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation

Qing Zhong, Peng-Tao Jiang, Wen Wang +3

Contemporary Video Instance Segmentation (VIS) methods typically adhere to a pre-train then fine-tune regime, where a segmentation model trained on images is fine-tuned on videos.…