180 citations · 184 across the 8 of their papers we have counts for
8 papers
QQQ: Quality Quattuor-Bit Quantization for Large Language Models
Ying Zhang, Peng Zhang, Mincong Huang +7
Quantization is a proven effective method for compressing large language models. Although popular techniques like W8A8 and W4A16 effectively maintain model performance, they often…
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Team GLM, :, Aohan Zeng +56
We introduce ChatGLM, an evolving family of large language models that we have been developing over time. This report primarily focuses on the GLM-4 language series, which includes…
On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models
Xinpeng Wang, Shitong Duan, Xiaoyuan Yi +7
Big models have achieved revolutionary breakthroughs in the field of AI, but they might also pose potential concerns. Addressing such concerns, alignment technologies were introduc…
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
Mincong Huang, Chao Wang, Chi Ma +3
Pipeline parallelism is an essential technique in the training of large-scale Transformer models. However, it suffers from imbalanced memory consumption, leading to insufficient me…
UGC: Unified GAN Compression for Efficient Image-to-Image Translation
Yuxi Ren, Jie Wu, Peng Zhang +6
Recent years have witnessed the prevailing progress of Generative Adversarial Networks (GANs) in image-to-image translation. However, the success of these GAN models hinges on pond…
3D Multiple Object Tracking on Autonomous Driving: A Literature Review
Peng Zhang, Xin Li, Liang He +1
3D multi-object tracking (3D MOT) stands as a pivotal domain within autonomous driving, experiencing a surge in scholarly interest and commercial promise over recent years. Despite…