6 papers
Weight Tying Biases Token Embeddings Towards the Output Space
Antonio Lopardo, Avyukth Harish, Catherine Arnett +1
Weight tying, i.e. sharing parameters between input and output embedding matrices, is common practice in language model design, yet its impact on the learned embedding space remain…
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
Cheol Jun Cho, Nicholas Lee, Akshat Gupta +4
Syllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such str…
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
Akshat Gupta, Atahan Ozdemir, Gopala Anumanchipalli
This paper presents a novel geometric interpretation of LayerNorm and explores how LayerNorm influences the norm and orientation of hidden vectors in the representation space. With…
PokerBench: Training Large Language Models to become Professional Poker Players
Richard Zhuang, Akshat Gupta, Richard Yang +3
We introduce PokerBench - a benchmark for evaluating the poker-playing abilities of large language models (LLMs). As LLMs excel in traditional NLP tasks, their application to compl…
A Unified Framework for Model Editing
Akshat Gupta, Dev Sajnani, Gopala Anumanchipalli
ROME and MEMIT are largely believed to be two different model editing algorithms, with the major difference between them being the ability to perform batched edits. In this paper,…
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
Akshat Gupta, Sidharth Baskaran, Gopala Anumanchipalli
Recent work using Rank-One Model Editing (ROME), a popular model editing method, has shown that there are certain facts that the algorithm is unable to edit without breaking the mo…