collaborators

7 papers

cs.CV2025

Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images

JiaKui Hu, Shanshan Zhao, Qing-Guo Chen +6

This paper presents Omni-View, which extends the unified multimodal understanding and generation to 3D scenes based on multiview images, exploring the principle that "generation fa…

cs.CV2025

Ovis2.5 Technical Report

Shiyin Lu, Yang Li, Yu Xia +39

We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…

cs.CV2025

Ovis-U1 Technical Report

Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9

In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Buildi…

cs.CV2025

High-quality Pseudo-labeling for Point Cloud Segmentation with Scene-level Annotation

Lunhao Duan, Shanshan Zhao, Xingxing Weng +2

This paper investigates indoor point cloud semantic segmentation under scene-level annotation, which is less explored compared to methods relying on sparse point-level labels. In t…

cs.CV2025

Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities

Shanshan Zhao, Xinjie Zhang, Jintao Guo +9

Recent years have seen remarkable progress in both multimodal understanding models and image generation models. Despite their respective successes, these two domains have evolved i…

cs.CV2025

JointTuner: Appearance-Motion Adaptive Joint Training for Customized Video Generation

Fangda Chen, Shanshan Zhao, Chuanfu Xu +1

Recent advancements in customized video generation have led to significant improvements in the simultaneous adaptation of appearance and motion. Typically, decoupling the appearanc…