papers

Publications (17)

cs.LG2024

Faster and Lighter LLMs: A Survey on Current Challenges and Way Forward

Arnav Chavan, Raghav Magazine, Shubham Kushwaha +2

Despite the impressive performance of LLMs, their widespread adoption faces challenges due to substantial computational and memory requirements during inference. Recent advancement…

cs.LG2026

S2D: Selective Spectral Decay for Quantization-Friendly Conditioning of Neural Activations

Arnav Chavan, Nahush Lele, Udbhav Bamba +3

Activation outliers in large-scale transformer models pose a fundamental challenge to model quantization, creating excessively large ranges that cause severe accuracy drops during…

cs.LG2023

One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning

Arnav Chavan, Zhuang Liu, Deepak Gupta +2

We present Generalized LoRA (GLoRA), an advanced approach for universal parameter-efficient fine-tuning tasks. Enhancing Low-Rank Adaptation (LoRA), GLoRA employs a generalized pro…

cs.CV2022

Transfer Learning Gaussian Anomaly Detection by Fine-tuning Representations

Oliver Rippel, Arnav Chavan, Chucai Lei +1

Current state-of-the-art anomaly detection (AD) methods exploit the powerful representations yielded by large-scale ImageNet training. However, catastrophic forgetting prevents the…

cs.CE2024

Parameter Efficient Fine-Tuning for Deep Learning-Based Full-Waveform Inversion

Koustav Ghosal, Abhranta Panigrahi, Arnav Chavan +2

Seismic full waveform inversion (FWI) has seen promising advancements through deep learning. Existing approaches typically focus on task-specific models trained and evaluated in is…

cs.CV2021

Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopy

Sharib Ali, Mariia Dmitrieva, Noha Ghatwary +34

The Endoscopy Computer Vision Challenge (EndoCV) is a crowd-sourcing initiative to address eminent problems in developing reliable computer aided detection and diagnosis endoscopy…

cs.LG2023

Rethinking Compression: Reduced Order Modelling of Latent Features in Large Language Models

Arnav Chavan, Nahush Lele, Deepak Gupta

Due to the substantial scale of Large Language Models (LLMs), the direct application of conventional compression methodologies proves impractical. The computational demands associa…

cs.CV2021

Rescaling CNN through Learnable Repetition of Network Parameters

Arnav Chavan, Udbhav Bamba, Rishabh Tiwari +1

Deeper and wider CNNs are known to provide improved performance for deep learning tasks. However, most such networks have poor performance gain per parameter increase. In this pape…

cs.CV2021

ChipNet: Budget-Aware Pruning with Heaviside Continuous Approximations

Rishabh Tiwari, Udbhav Bamba, Arnav Chavan +1

Structured pruning methods are among the effective strategies for extracting small resource-efficient convolutional neural networks from their dense counterparts with minimal loss…

cs.CV2022

Vision Transformer Slimming: Multi-Dimension Searching in Continuous Optimization Space

Arnav Chavan, Zhiqiang Shen, Zhuang Liu +3

This paper explores the feasibility of finding an optimal sub-model from a vision transformer and introduces a pure vision transformer slimming (ViT-Slim) framework. It can search…

cs.LG2024

Beyond Uniform Scaling: Exploring Depth Heterogeneity in Neural Architectures

Akash Guna R. T, Arnav Chavan, Deepak Gupta

Conventional scaling of neural networks typically involves designing a base network and growing different dimensions like width, depth, etc. of the same by some predefined scaling…

cs.CV2023

On Designing Light-Weight Object Trackers through Network Pruning: Use CNNs or Transformers?

Saksham Aggarwal, Taneesh Gupta, Pawan Kumar Sahu +4

Object trackers deployed on low-power devices need to be light-weight, however, most of the current state-of-the-art (SOTA) methods rely on using compute-heavy backbones built usin…

cs.CV2023

Patch Gradient Descent: Training Neural Networks on Very Large Images

Deepak K. Gupta, Gowreesh Mago, Arnav Chavan +1

Traditional CNN models are trained and tested on relatively low resolution images (<300 px), and cannot be directly operated on large-scale images due to compute and memory constra…

cs.CV2020

Multi-Plateau Ensemble for Endoscopic Artefact Segmentation and Detection

Suyog Jadhav, Udbhav Bamba, Arnav Chavan +2

Endoscopic artefact detection challenge consists of 1) Artefact detection, 2) Semantic segmentation, and 3) Out-of-sample generalisation. For Semantic segmentation task, we propose…

cs.LG2022

Dynamic Kernel Selection for Improved Generalization and Memory Efficiency in Meta-learning

Arnav Chavan, Rishabh Tiwari, Udbhav Bamba +1

Gradient based meta-learning methods are prone to overfit on the meta-training set, and this behaviour is more prominent with large and complex networks. Moreover, large networks r…

cs.LG2026

DOT-MoE: Differentiable Optimal Transport for MoEfication

Udbhav Bamba, Arnav Chavan, Aryamaan Thakur +2

The scaling of Large Language Models (LLMs) has driven significant performance gains but created substantial challenges in inference efficiency. While Mixture of Experts (MoEs) arc…

cs.CL2024

Surgical Feature-Space Decomposition of LLMs: Why, When and How?

Arnav Chavan, Nahush Lele, Deepak Gupta

Low-rank approximations, of the weight and feature space can enhance the performance of deep learning models, whether in terms of improving generalization or reducing the latency o…