6 papers
Ovis2.5 Technical Report
Shiyin Lu, Yang Li, Yu Xia +39
We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning
Yiyang Chen, Shanshan Zhao, Lunhao Duan +2
Diffusion-based models, widely used in text-to-image generation, have proven effective in 2D representation learning. Recently, this framework has been extended to 3D self-supervis…
Ovis-U1 Technical Report
Guo-Hua Wang, Shanshan Zhao, Xinjie Zhang +9
In this report, we introduce Ovis-U1, a 3-billion-parameter unified model that integrates multimodal understanding, text-to-image generation, and image editing capabilities. Buildi…
High-quality Pseudo-labeling for Point Cloud Segmentation with Scene-level Annotation
Lunhao Duan, Shanshan Zhao, Xingxing Weng +2
This paper investigates indoor point cloud semantic segmentation under scene-level annotation, which is less explored compared to methods relying on sparse point-level labels. In t…
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
Shanshan Zhao, Xinjie Zhang, Jintao Guo +9
Recent years have seen remarkable progress in both multimodal understanding models and image generation models. Despite their respective successes, these two domains have evolved i…
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
Lunhao Duan, Shanshan Zhao, Wenjun Yan +7
Recently, text-to-image generation models have achieved remarkable advancements, particularly with diffusion models facilitating high-quality image synthesis from textual descripti…