3 papers
cs.AI2026
F-WANDA: Fisher-Reweighted Post-Training Pruning for Sustainable Deployment of Large Language Models
Himanshu Mishra
One-shot post-training pruning is the most energy-frugal compression strategy for largelanguage models (LLMs), yet existing approaches trade either quality (WANDA) or compute cost…
cs.LG2026
RDumb++: Drift-Aware Continual Test-Time Adaptation
Himanshu Mishra
Continual Test-Time Adaptation (CTTA) seeks to update a pretrained model during deployment using only the incoming, unlabeled data stream. Although prior approaches such as Tent, E…
cs.LG2026
QUAIL: Quantization Aware Unlearning for Mitigating Misinformation in LLMs
Himanshu Mishra, Kanwal Mehreen
Machine unlearning aims to remove specific knowledge (e.g., copyrighted or private data) from a trained model without full retraining. In practice, models are often quantized (e.g.…