activity
20242026
collaborators

13 papers

cs.CV2026

FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning

Bo Yin, Xiaobin Hu, Xingyu Zhou +7

Diffusion models have achieved remarkable success in generative modeling, yet how to effectively adapt large pretrained models to new tasks remains challenging. We revisit the reco…

cs.CV2026

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

Yinan Chen, Chuming Lin, Zhennan Chen +12

While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets and benchmarks. To bridge t…

cs.CV2026

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset

Haojun Chen, Haoyang He, Chengming Xu +11

Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the rapid development of imagin…

cs.CV2026

DiP: Taming Diffusion Models in Pixel Space

Zhennan Chen, Junwei Zhu, Xu Chen +6

Diffusion models face a fundamental trade-off between generation quality and computational efficiency. Latent Diffusion Models (LDMs) offer an efficient solution but suffer from po…

cs.CV2025

Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation

Jiangning Zhang, Junwei Zhu, Zhenye Gan +14

We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed , which generates semantically coherent videos from a single-fram…

cs.CV2025

Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10

Jiangning Zhang, Junwei Zhu, Teng Hu +7

Native 4K (21603840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, ma…