Publications (22)
A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
Jiacheng Liu, Xinyu Wang, Yuqi Lin +10
Diffusion Models have become a cornerstone of modern generative AI for their exceptional generation quality and controllability. However, their inherent \textit{multi-step iteratio…
CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation
Yuqi Lin, Minghao Chen, Wenxiao Wang +5
Weakly supervised semantic segmentation (WSSS) with image-level labels is a challenging task. Mainstream approaches follow a multi-stage framework and suffer from high training cos…
HiCache: A Plug-in Scaled-Hermite Upgrade for Taylor-Style Cache-then-Forecast Diffusion Acceleration
Liang Feng, Shikang Zheng, Jiacheng Liu +8
Diffusion models have achieved remarkable success in content generation but often incur prohibitive computational costs due to iterative sampling. Recent feature caching methods ac…
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
Yuqi Lin, Hengjia Li, Wenqi Shao +5
In this paper, we explore a principal way to enhance the quality of widely pre-existing coarse masks, enabling them to serve as reliable training data for segmentation models to re…
Self-supervised and Weakly Supervised Contrastive Learning for Frame-wise Action Representations
Minghao Chen, Renbo Tu, Chenxi Huang +3
Previous work on action representation learning focused on global representations for short video clips. In contrast, many practical applications, such as video alignment, strongly…
From Sketch to Fresco: Efficient Diffusion Transformer with Progressive Resolution
Shikang Zheng, Guantao Chen, Lixuan He +4
Diffusion Transformers achieve impressive generative quality but remain computationally expensive due to iterative sampling. Recently, dynamic resolution sampling has emerged as a…
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
Pengfei Zhou, Xiaopeng Peng, Jiajun Song +15
Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a ch…
Few-shot Hybrid Domain Adaptation of Image Generators
Hengjia Li, Yang Liu, Linxuan Xia +7
Can a pre-trained generator be adapted to the hybrid of multiple target domains and generate images with integrated attributes of them? In this work, we introduce a new task -- Few…
UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation
Hengjia Li, Yang Liu, Yuqi Lin +8
Recently, generative domain adaptation has achieved remarkable progress, enabling us to adapt a pre-trained generator to a new target domain. However, existing methods simply adapt…
Forecast then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers
Shikang Zheng, Liang Feng, Xinyu Wang +8
Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To reduce their substantial computational costs, feature cachin…
Position: Towards Implicit Prompt For Text-To-Image Models
Yue Yang, Yuqi Lin, Hong Liu +7
Recent text-to-image (T2I) models have had great success, and many benchmarks have been proposed to evaluate their performance and safety. However, they only consider explicit prom…
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
Shuo Liu, Kaining Ying, Hao Zhang +8
This paper presents ConvBench, a novel multi-turn conversation evaluation benchmark tailored for Large Vision-Language Models (LVLMs). Unlike existing benchmarks that assess indivi…
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Kaining Ying, Fanqing Meng, Jin Wang +19
Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimod…
Free-T2M: Robust Text-to-Motion Generation for Humanoid Robots via Frequency-Domain
Wenshuo Chen, Haozhe Jia, Songning Lai +5
Enabling humanoid robots to synthesize complex, physically coherent motions from natural language commands is a cornerstone of autonomous robotics and human-robot interaction. Whil…
FreqCa: Accelerating Diffusion Models via Frequency-Aware Caching
Jiacheng Liu, Peiliang Cai, Qinming Zhou +9
The application of diffusion transformers is suffering from their significant inference costs. Recently, feature caching has been proposed to solve this problem by reusing features…
TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP Without Training
Yuqi Lin, Minghao Chen, Kaipeng Zhang +7
Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in open-vocabulary classification. The class token in the image encoder is trained to captur…
Dynamic Video Generation: Shaping Video Generation Across Time and Space
Shikang Zheng, Jingkai Huang, Jiacheng Liu +5
Diffusion models have achieved impressive performance in video generation, but their iterative denoising process remains computationally expensive due to the large number of tokens…
Cross-sphere Coupling and Source Inversion of Ionospheric Disturbances Associated with the 2025 Myanmar Strike-slip Earthquake from BeiDou GEO and Multi-GNSS Observations
Jianghe Chen, Pan Xiong, Qingshan Ruan +6
Focusing on the M7.9 earthquake in Myanmar in 2025, this study comprehensively utilizes data from BeiDou geostationary satellites of the Chinese Continental Crustal Movement Observ…
Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Efficient Diffusion Transformers
Guantao Chen, Shikang Zheng, Yuqi Lin +1
Diffusion Transformer (DiT) models have achieved unprecedented quality in image and video generation, yet their iterative sampling process remains computationally prohibitive. To a…
Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
Shikang Zheng, Guantao Chen, Qinming Zhou +6
Diffusion Transformers offer state-of-the-art fidelity in image and video synthesis, but their iterative sampling process remains a major bottleneck due to the high cost of transfo…
SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking
Zhengan Yan, Shikang Zheng, Haoran Qin +9
Diffusion-based image editing offers strong semantic controllability, but remains computationally expensive due to iterative high-resolution denoising over all spatial tokens. Dyna…
LUMA: Low-Dimension Unified Motion Alignment with Dual-Path Anchoring for Text-to-Motion Diffusion Model
Haozhe Jia, Wenshuo Chen, Yuqi Lin +8
While current diffusion-based models, typically built on U-Net architectures, have shown promising results on the text-to-motion generation task, they still suffer from semantic mi…