5 papers
Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing
Lin Liu, Zhihan Xiao, Haohang Xu +4
Video editing has recently achieved remarkable progress with diffusion-based generative models, enabling diverse object-level manipulations from natural language instructions. Howe…
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
Haohang Xu, Lin Liu, Zhibo Zhang +3
Diffusion-based image editing models have achieved significant progress in real world applications. However, conventional models typically rely on natural language prompts, which o…
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
Peisen Zhao, Xiaopeng Zhang, Mingxing Xu +10
While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. Conventional CLIP-based encode…
ELiTe: Efficient Image-to-LiDAR Knowledge Transfer for Semantic Segmentation
Zhibo Zhang, Ximing Yang, Weizhong Zhang +1
Cross-modal knowledge transfer enhances point cloud representation learning in LiDAR semantic segmentation. Despite its potential, the \textit{weak teacher challenge} arises due to…
Robust Fine-tuning for Pre-trained 3D Point Cloud Models
Zhibo Zhang, Ximing Yang, Weizhong Zhang +1
This paper presents a robust fine-tuning method designed for pre-trained 3D point cloud models, to enhance feature robustness in downstream fine-tuned models. We highlight the limi…