1 citations · 1 across the 3 of their papers we have counts for
4 papers
Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation
Tianyi Zeng, Junchao Liao, Yujie Wei +9
Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherence does not imply the ability to execute editing techniques. Profession…
Ovis2.5 Technical Report
Shiyin Lu, Yang Li, Yu Xia +39
We present Ovis2.5, a successor to Ovis2 designed for native-resolution visual perception and strong multimodal reasoning. Ovis2.5 integrates a native-resolution vision transformer…
PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm
Tianyu Chang, Xiaohao Chen, Zhichao Wei +5
Video Virtual Try-on aims to seamlessly transfer a reference garment onto a target person in a video while preserving both visual fidelity and temporal coherence. Existing methods…
MM-Diff: High-Fidelity Image Personalization via Multi-Modal Condition Integration
Zhichao Wei, Qingkun Su, Long Qin +1
Recent advances in tuning-free personalized image generation based on diffusion models are impressive. However, to improve subject fidelity, existing methods either retrain the dif…