2 papers
cs.LG2025
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
Dev Patel, Gabrielle Gervacio, Diekola Raimi +5
Large Language Models require substantial computational resources for inference, posing deployment challenges. While dynamic pruning offers superior efficiency over static methods…
cs.LG2025
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
Siddharth Chaudhary, Dev Patel, Maheep Chaudhary +1
The quadratic complexity of transformers fundamentally limits reasoning system deployment in resource-constrained and long-context settings. We introduce Hydra, a modular architect…