6 papers
LayerT2V: A Unified Multi-Layer Video Generation Framework
Guangzhao Li, Kangrui Cen, Baixuan Zhao +5
Text-to-video generation has advanced rapidly, but existing methods typically output only the final composited video and lack editable layered representations, limiting their use i…
TR-PTS: Task-Relevant Parameter and Token Selection for Efficient Tuning
Siqi Luo, Haoran Yang, Yi Xin +4
Large pre-trained models achieve remarkable performance in vision tasks but are impractical for fine-tuning due to high computational and storage costs. Parameter-Efficient Fine-Tu…
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
Yi Xin, Juncheng Yan, Qi Qin +18
We present Lumina-mGPT 2.0, a stand-alone, decoder-only autoregressive model that revisits and revitalizes the autoregressive paradigm for high-quality image generation and beyond.…
Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
Yi Xin, Le Zhuo, Qi Qin +8
AutoRegressive (AR) models have made notable progress in image generation, with Masked AutoRegressive (MAR) models gaining attention for their efficient parallel decoding. However,…
Towards Effective User Attribution for Latent Diffusion Models via Watermark-Informed Blending
Yongyang Pan, Xiaohong Liu, Siqi Luo +5
Rapid advancements in multimodal large language models have enabled the creation of hyper-realistic images from textual descriptions. However, these advancements also raise signifi…
Enhancing Test Time Adaptation with Few-shot Guidance
Siqi Luo, Yi Xin, Yuntao Du +3
Deep neural networks often encounter significant performance drops while facing with domain shifts between training (source) and test (target) data. To address this issue, Test Tim…