6 papers
Learnable Permutation for Structured Sparsity on Transformer Models
Zekai Li, Ji Liu, Guanchen Li +5
Structured sparsity has emerged as a popular model pruning technique, widely adopted in various architectures, including CNNs, Transformer models, and especially large language mod…
DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code Generation
Jiajun jiao, Haowei Zhu, Puyuan Yang +8
Diffusion models have achieved remarkable success in image and video generation. However, their inherently multiple step inference process imposes substantial computational overhea…
Progressive Conditioned Scale-Shift Recalibration of Self-Attention for Online Test-time Adaptation
Yushun Tang, Ziqiong Liu, Jiyuan Jia +2
Online test-time adaptation aims to dynamically adjust a network model in real-time based on sequential input samples during the inference stage. In this work, we find that, when a…
Open-World Test-Time Adaptation with Hierarchical Feature Aggregation and Attention Affine
Ziqiong Liu, Yushun Tang, Junyang Ji +1
Test-time adaptation (TTA) refers to adjusting the model during the testing phase to cope with changes in sample distribution and enhance the model's adaptability to new environmen…
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
Jianghui Wang, Vinay Joshi, Saptarshi Majumder +7
The demand for AI-generated GPU kernels is rapidly growing, influenced by the need for scalable, hardware-optimized solutions in both industry and academia. As deep learning worklo…
Visual Reranking with Improved Image Graph
Ziqiong Liu, Shengjin Wang, Liang Zheng +1
This paper introduces an improved reranking method for the Bag-of-Words (BoW) based image search. Built on [1], a directed image graph robust to outlier distraction is proposed. In…