4 papers
BeCARE: Budgeted Cache Refresh for Diffusion Transformer Acceleration
Yuhang Zhang, Junxiang Qiu, Huixia Ben +4
Training-free feature caching accelerates diffusion transformer (DiT) inference by reusing or forecasting intermediate features. However, fixed schedules make compute predictable b…
Image Captioning via Compact Bidirectional Architecture
Zijie Song, Yuanen Zhou, Zhenzhen Hu +4
Most current image captioning models typically generate captions from left-to-right. This unidirectional property makes them can only leverage past context but not future context.…
Accelerating Controllable Generation via Hybrid-grained Cache
Lin Liu, Huixia Ben, Shuo Wang +4
Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation…
StgcDiff: Spatial-Temporal Graph Condition Diffusion for Sign Language Transition Generation
Jiashu He, Jiayi He, Shengeng Tang +3
Sign language transition generation seeks to convert discrete sign language segments into continuous sign videos by synthesizing smooth transitions. However,most existing methods m…