3 papers
cs.CL2025
Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing
James O' Neill, Santhosh Subramanian, Eric Lin +1
The trend towards large language models (LLMs) for guardrailing against undesired behaviors is increasing and has shown promise for censoring user inputs. However, increased latenc…
cs.CL2023
Gradient Sparsification For Masked Fine-Tuning of Transformers
James O' Neill, Sourav Dutta
Fine-tuning pretrained self-supervised language models is widely adopted for transfer learning to downstream tasks. Fine-tuning can be achieved by freezing gradients of the pretrai…
cs.CL2023
Self-Distilled Quantization: Achieving High Compression Rates in Transformer-Based Language Models
James O' Neill, Sourav Dutta
We investigate the effects of post-training quantization and quantization-aware training on the generalization of Transformer language models. We present a new method called self-d…