collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

Dynamic Short Convolutions Improve Transformers

Oliver Sieberling, Bharat Runwal, Rameswar Panda +1

Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forward layers, residual connections…

cs.LG2026

PRISM: Demystifying Retention and Interaction in Mid-Training

Bharat Runwal, Ashish Agrawal, Anurag Roy +1

We present PRISM, a comprehensive empirical study of mid-training design choices for large language models. Through controlled experiments across seven base models spanning four fa…

cs.LG2024

SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning

Jinghan Jia, Yihua Zhang, Yimeng Zhang +5

Large Language Models (LLMs) have highlighted the necessity of effective unlearning mechanisms to comply with data regulations and ethical AI practices. LLM unlearning aims at remo…

cs.LG2024

From PEFT to DEFT: Parameter Efficient Finetuning for Reducing Activation Density in Transformers

Bharat Runwal, Tejaswini Pedapati, Pin-Yu Chen

Pretrained Language Models (PLMs) have become the de facto starting point for fine-tuning on downstream tasks. However, as model sizes continue to increase, traditional fine-tuning…

cs.LG2023

Uncovering the Hidden Cost of Model Compression

Diganta Misra, Muawiz Chaudhary, Agam Goyal +2

In an age dominated by resource-intensive foundation models, the ability to efficiently adapt to downstream tasks is crucial. Visual Prompting (VP), drawing inspiration from the pr…