3 papers
cs.LG2026
EdgeRazor: A Lightweight Framework for Large Language Models via Mixed-Precision Quantization-Aware Distillation
Shu-Hao Zhang, Le-Tong Huang, Xiang-Sheng Deng +5
Quantization has emerged as a mainstream approach for deploying Large Language Models (LLMs) on resource-constrained devices, yet compressing precision below 4-bit typically causes…
cs.CV2025
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
Shu-Hao Zhang, Wei-Cheng Tang, Chen Wu +5
Recent years have witnessed an increasing interest in image-text contrastive modeling, exemplified by models such as Contrastive Language-Image Pretraining (CLIP). In this paper, w…
cs.CL2024
Efficient Ternary Weight Embedding Model: Bridging Scalability and Performance
Jiayi Chen, Chen Wu, Shaoqun Zhang +3
Embedding models have become essential tools in both natural language processing and computer vision, enabling efficient semantic search, recommendation, clustering, and more. Howe…