collaborators

5 papers

cs.CV2026

Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing

Lin Liu, Zhihan Xiao, Haohang Xu +4

Video editing has recently achieved remarkable progress with diffusion-based generative models, enabling diverse object-level manipulations from natural language instructions. Howe…

cs.CV2026

FineEdit: Fine-Grained Image Edit with Bounding Box Guidance

Haohang Xu, Lin Liu, Zhibo Zhang +3

Diffusion-based image editing models have achieved significant progress in real world applications. However, conventional models typically rely on natural language prompts, which o…

cs.CV2026

FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions

Peisen Zhao, Xiaopeng Zhang, Mingxing Xu +10

While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. Conventional CLIP-based encode…

cs.CV2024

ELiTe: Efficient Image-to-LiDAR Knowledge Transfer for Semantic Segmentation

Zhibo Zhang, Ximing Yang, Weizhong Zhang +1

Cross-modal knowledge transfer enhances point cloud representation learning in LiDAR semantic segmentation. Despite its potential, the \textit{weak teacher challenge} arises due to…

cs.CV2024

Robust Fine-tuning for Pre-trained 3D Point Cloud Models

Zhibo Zhang, Ximing Yang, Weizhong Zhang +1

This paper presents a robust fine-tuning method designed for pre-trained 3D point cloud models, to enhance feature robustness in downstream fine-tuned models. We highlight the limi…