1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2025
TuneComp: Joint Fine-tuning and Compression for Large Foundation Models
Xiangyu Chen, Jing Liu, Ye Wang +4
To reduce model size during post-training, compression methods, including knowledge distillation, low-rank approximation, and pruning, are often applied after fine-tuning the model…
cs.LG2025
LatentLLM: Attention-Aware Joint Tensor Compression
Toshiaki Koike-Akino, Xiangyu Chen, Jing Liu +4
Modern foundation models such as large language models (LLMs) and large multi-modal models (LMMs) require a massive amount of computational and memory resources. We propose a new f…
cs.CV2024★ 1 cited
SuperLoRA: Parameter-Efficient Unified Adaptation of Multi-Layer Attention Modules
Xiangyu Chen, Jing Liu, Ye Wang +4
Low-rank adaptation (LoRA) and its variants are widely employed in fine-tuning large models, including large language models for natural language processing and diffusion models fo…