papers

Publications (22)

cs.LG2025

A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation

Jiacheng Liu, Xinyu Wang, Yuqi Lin +10

Diffusion Models have become a cornerstone of modern generative AI for their exceptional generation quality and controllability. However, their inherent \textit{multi-step iteratio…

cs.CV2023

CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation

Yuqi Lin, Minghao Chen, Wenxiao Wang +5

Weakly supervised semantic segmentation (WSSS) with image-level labels is a challenging task. Mainstream approaches follow a multi-stage framework and suffer from high training cos…

cs.CV2026

HiCache: A Plug-in Scaled-Hermite Upgrade for Taylor-Style Cache-then-Forecast Diffusion Acceleration

Liang Feng, Shikang Zheng, Jiacheng Liu +8

Diffusion models have achieved remarkable success in content generation but often incur prohibitive computational costs due to iterative sampling. Recent feature caching methods ac…

cs.CV2025

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement

Yuqi Lin, Hengjia Li, Wenqi Shao +5

In this paper, we explore a principal way to enhance the quality of widely pre-existing coarse masks, enabling them to serve as reliable training data for segmentation models to re…

cs.CV2023

Self-supervised and Weakly Supervised Contrastive Learning for Frame-wise Action Representations

Minghao Chen, Renbo Tu, Chenxi Huang +3

Previous work on action representation learning focused on global representations for short video clips. In contrast, many practical applications, such as video alignment, strongly…

cs.CV2026

From Sketch to Fresco: Efficient Diffusion Transformer with Progressive Resolution

Shikang Zheng, Guantao Chen, Lixuan He +4

Diffusion Transformers achieve impressive generative quality but remain computationally expensive due to iterative sampling. Recently, dynamic resolution sampling has emerged as a…

cs.CV2025

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

Pengfei Zhou, Xiaopeng Peng, Jiajun Song +15

Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a ch…

cs.CV2023

Few-shot Hybrid Domain Adaptation of Image Generators

Hengjia Li, Yang Liu, Linxuan Xia +7

Can a pre-trained generator be adapted to the hybrid of multiple target domains and generate images with integrated attributes of them? In this work, we introduce a new task -- Few…

cs.CV2024

UniHDA: A Unified and Versatile Framework for Multi-Modal Hybrid Domain Adaptation

Hengjia Li, Yang Liu, Yuqi Lin +8

Recently, generative domain adaptation has achieved remarkable progress, enabling us to adapt a pre-trained generator to a new target domain. However, existing methods simply adapt…

cs.CV2025

Forecast then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers

Shikang Zheng, Liang Feng, Xinyu Wang +8

Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To reduce their substantial computational costs, feature cachin…

cs.CY2024

Position: Towards Implicit Prompt For Text-To-Image Models

Yue Yang, Yuqi Lin, Hong Liu +7

Recent text-to-image (T2I) models have had great success, and many benchmarks have been proposed to evaluate their performance and safety. However, they only consider explicit prom…

cs.MM2024

ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models

Shuo Liu, Kaining Ying, Hao Zhang +8

This paper presents ConvBench, a novel multi-turn conversation evaluation benchmark tailored for Large Vision-Language Models (LVLMs). Unlike existing benchmarks that assess indivi…

cs.CV2024

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Kaining Ying, Fanqing Meng, Jin Wang +19

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimod…

cs.CV2025

Free-T2M: Robust Text-to-Motion Generation for Humanoid Robots via Frequency-Domain

Wenshuo Chen, Haozhe Jia, Songning Lai +5

Enabling humanoid robots to synthesize complex, physically coherent motions from natural language commands is a cornerstone of autonomous robotics and human-robot interaction. Whil…

cs.LG2025

FreqCa: Accelerating Diffusion Models via Frequency-Aware Caching

Jiacheng Liu, Peiliang Cai, Qinming Zhou +9

The application of diffusion transformers is suffering from their significant inference costs. Recently, feature caching has been proposed to solve this problem by reusing features…

cs.CV2023

TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP Without Training

Yuqi Lin, Minghao Chen, Kaipeng Zhang +7

Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in open-vocabulary classification. The class token in the image encoder is trained to captur…

cs.CV2026

Dynamic Video Generation: Shaping Video Generation Across Time and Space

Shikang Zheng, Jingkai Huang, Jiacheng Liu +5

Diffusion models have achieved impressive performance in video generation, but their iterative denoising process remains computationally expensive due to the large number of tokens…

physics.geo-ph2026

Cross-sphere Coupling and Source Inversion of Ionospheric Disturbances Associated with the 2025 Myanmar Strike-slip Earthquake from BeiDou GEO and Multi-GNSS Observations

Jianghe Chen, Pan Xiong, Qingshan Ruan +6

Focusing on the M7.9 earthquake in Myanmar in 2025, this study comprehensively utilizes data from BeiDou geostationary satellites of the Chinese Continental Crustal Movement Observ…

cs.CV2026

Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Efficient Diffusion Transformers

Guantao Chen, Shikang Zheng, Yuqi Lin +1

Diffusion Transformer (DiT) models have achieved unprecedented quality in image and video generation, yet their iterative sampling process remains computationally prohibitive. To a…

cs.CV2025

Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers

Shikang Zheng, Guantao Chen, Qinming Zhou +6

Diffusion Transformers offer state-of-the-art fidelity in image and video synthesis, but their iterative sampling process remains a major bottleneck due to the high cost of transfo…

cs.CV2026

SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking

Zhengan Yan, Shikang Zheng, Haoran Qin +9

Diffusion-based image editing offers strong semantic controllability, but remains computationally expensive due to iterative high-resolution denoising over all spatial tokens. Dyna…

cs.CV2025

LUMA: Low-Dimension Unified Motion Alignment with Dual-Path Anchoring for Text-to-Motion Diffusion Model

Haozhe Jia, Wenshuo Chen, Yuqi Lin +8

While current diffusion-based models, typically built on U-Net architectures, have shown promising results on the text-to-motion generation task, they still suffer from semantic mi…