activity
20242026
collaborators

5 papers

cs.CL2026

There is No Theoretical Curse of Multilinguality For Embedding Space Structure

Niyati Bafna, Neha Verma, Vilém Zouhar +2

A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model.…

cs.CL2026

ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging

Neha Verma, Nikhil Mehta, Shao-Chuan Wang +7

Despite the rapid advancements in large language model (LLM) development, fine-tuning them for specific tasks often results in the catastrophic forgetting of their general, languag…

cs.LG2026

DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging

Neha Verma, Kenton Murray, Kevin Duh

Structured pruning methods designed for Large Language Models (LLMs) generally focus on identifying and removing the least important components to optimize model size. However, in…

cs.CL2025

Merging Feed-Forward Sublayers for Compressed Transformers

Neha Verma, Kenton Murray, Kevin Duh

With the rise and ubiquity of larger deep learning models, the need for high-quality compression techniques is growing in order to deploy these models widely. The sheer parameter c…

cs.CL2024

Merging Text Transformer Models from Different Initializations

Neha Verma, Maha Elbayad

Recent work on permutation-based model merging has shown impressive low- or zero-barrier mode connectivity between models from completely different initializations. However, this l…