4 papers
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
Jie Kong, Wei Wang, Jiehan Zhou +1
Major challenges in LLMs inference remain frequent memory bandwidth bottlenecks, computational redundancy, and inefficiencies in long-sequence processing. To address these issues,…
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
Yaozheng Zhang, Wei Wang, Jie Kong +5
The increasing adoption of large language models (LLMs) on heterogeneous computing platforms poses significant challenges to achieving high inference efficiency. To address these e…
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
Jie Kong, Junxiang Zhang, Jiheng Xu +7
In the field of deep learning, traditional attention mechanisms face significant challenges related to high computational complexity and large memory consumption when processing lo…
InsightEdit: Towards Better Instruction Following for Image Editing
Yingjing Xu, Jie Kong, Jiazhi Wang +3
In this paper, we focus on the task of instruction-based image editing. Previous works like InstructPix2Pix, InstructDiffusion, and SmartEdit have explored end-to-end editing. Howe…