14 papers
COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models
Peiqi Yu, Nam Ling, Wei Wang +1
Structured pruning reduces the size and inference cost of large language models (LLMs) by removing weight columns, but the resulting output error can degrade accuracy. Existing tra…
Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies
Wei Jiang, Wei Wang
Vision-Language-Action (VLA) models process long multimodal token sequences, making inference expensive in both memory and computation. Existing efficiency methods mainly reduce vi…
CORAM: Coherent Orthogonal Rotation for Model Merging
Xinyi Sui, Ziran Liu, Nam Ling +2
Merging finetuned models combines specialized capabilities without joint training or access to the original data. Most methods operate by linear arithmetic in Euclidean weight spac…
CORA: Per-Slice Coherent Orthogonal Rotation for SVD-based Low-Rank Adaptation
Pengcheng Wang, Ziran Liu, Wei Wang +1
Parameter-Efficient Fine-Tuning (PEFT) commonly adapts pretrained weights through low-rank updates, and recent methods further exploit the singular value decomposition (SVD) of the…
Sub-Token Routing for KV Cache Compression
Wei Jiang, Wei Wang
Transformer inference often requires a large KV cache, especially for long-context language modeling and multimodal generation. Existing compression methods usually reduce cache co…
Geometric and Spectral Alignment for Deep Neural Network II
Ziran Liu, Wei Wang, Jinhao Wang +5
This paper develops the angular and static-channel component of Geometric and Spectral Alignment for residual Jacobian chains. Starting from Cartan-coordinate rigidity and fitted e…