1 citations · 2 across the 5 of their papers we have counts for
4 papers · 1 filter
Ovis2.5 Technical Report
Shiyin Lu, Yang Li, Yu Xia +39
We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…
PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm
Tianyu Chang, Xiaohao Chen, Zhichao Wei +5
Video Virtual Try-on aims to seamlessly transfer a reference garment onto a target person in a video while preserving both visual fidelity and temporal coherence. Existing methods…
MM-Diff: High-Fidelity Image Personalization via Multi-Modal Condition Integration
Zhichao Wei, Qingkun Su, Long Qin +1
Recent advances in tuning-free personalized image generation based on diffusion models are impressive. However, to improve subject fidelity, existing methods either retrain the dif…
Linguistic Query-Guided Mask Generation for Referring Image Segmentation
Zhichao Wei, Xiaohao Chen, Mingqiang Chen +1
Referring image segmentation aims to segment the image region of interest according to the given language expression, which is a typical multi-modal task. Existing methods either a…