5 papers
Efficient Temporal-aware Matryoshka Adaptation for Temporal Information Retrieval
Tuan-Luc Huynh, Weiqing Wang, Trung Le +4
Retrievers are a key bottleneck in Temporal Retrieval-Augmented Generation (RAG) systems: failing to retrieve temporally relevant context can degrade downstream generation, regardl…
Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
Cuong Pham, Hoang Anh Dung, Cuong C. Nguyen +4
Large language models require significant computational resources for deployment, making quantization essential for practical applications. However, the main obstacle to effective…
Layer-Wise High-Impact Parameter Ratio Optimization in Post-Training Quantization for Large Language Models
Cuong Pham, Hoang Anh Dung, Cuong C. Nguyen +3
Large language models (LLMs) have significantly advanced natural language processing, but their massive parameter counts create substantial computational and memory challenges duri…
Optimizing Specific and Shared Parameters for Efficient Parameter Tuning
Van-Anh Nguyen, Thanh-Toan Do, Mehrtash Harandi +2
Foundation models, with a vast number of parameters and pretraining on massive datasets, achieve state-of-the-art performance across various applications. However, efficiently adap…
Enhancing Dataset Distillation via Non-Critical Region Refinement
Minh-Tuan Tran, Trung Le, Xuan-May Le +2
Dataset distillation has become a popular method for compressing large datasets into smaller, more efficient representations while preserving critical information for model trainin…