11 papers
Rethinking Output Alignment For 1-bit Post-Training Quantization of Large Language Models
Dung Anh Hoang, Cuong Pham, Cuong Nguyen +3
Large Language Models (LLMs) deliver strong performance across a wide range of NLP tasks, but their massive sizes hinder deployment on resource-constrained devices. To reduce their…
DWA-KD: Dual-Space Weighting and Time-Warped Alignment for Cross-Tokenizer Knowledge Distillation
Duc Trung Vu, Pham Khanh Chi, Dat Phi Van +3
Knowledge Distillation (KD) has emerged as a crucial technique for compressing Large Language Models (LLMs). Although existing cross-tokenizer KD methods have made notable progress…
Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment
Anh Bui, Trang Vu, Trung Le +5
In this paper, we investigate the semantic collapsing problem in generative personalization, an under-explored topic where the learned visual concept () gradually shifts from it…
On the Mechanisms of Collaborative Learning in VAE Recommenders
Tung-Long Vuong, Julien Monteil, Hien Dang +3
Variational Autoencoders (VAEs) are a powerful alternative to matrix factorization for recommendation. A common technique in VAE-based collaborative filtering (CF) consists in appl…
Efficient Temporal-aware Matryoshka Adaptation for Temporal Information Retrieval
Tuan-Luc Huynh, Weiqing Wang, Trung Le +4
Retrievers are a key bottleneck in Temporal Retrieval-Augmented Generation (RAG) systems: failing to retrieve temporally relevant context can degrade downstream generation, regardl…
Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
Cuong Pham, Hoang Anh Dung, Cuong C. Nguyen +4
Large language models require significant computational resources for deployment, making quantization essential for practical applications. However, the main obstacle to effective…