activity
20242026
collaborators

5 papers

cs.CV2026

DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval

Xinwei He, Yansong Zheng, Qianru Han +7

Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent…

cs.CV2025

Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation

Xin Yan, Yuxuan Cai, Qiuyue Wang +3

We introduce Presto, a novel video diffusion model designed to generate 15-second videos with long-range coherence and rich content. Extending video generation methods to maintain…

cs.CV2025

Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention

Huiguo He, Qiuyue Wang, Yuan Zhou +4

Training-free diffusion models have achieved remarkable progress in generating multi-subject consistent images within open-domain scenarios. The key idea of these methods is to inc…

cs.CV2024

Fleximo: Towards Flexible Text-to-Human Motion Video Generation

Yuhang Zhang, Yuan Zhou, Zeyu Liu +4

Current methods for generating human motion videos rely on extracting pose sequences from reference videos, which restricts flexibility and control. Additionally, due to the limita…

cs.CV2024

Allegro: Open the Black Box of Commercial-Level Video Generation Model

Yuan Zhou, Qiuyue Wang, Yuxuan Cai +1

Significant advancements have been made in the field of video generation, with the open-source community contributing a wealth of research papers and tools for training high-qualit…