6 papers
Stronger Normalization-Free Transformers
Mingzhi Chen, Taiming Lu, Jiachen Zhu +2
Although normalization layers have long been viewed as indispensable components of deep learning architectures, the recent introduction of Dynamic Tanh (DyT) has demonstrated that…
Command-V: Pasting LLM Behaviors via Activation Profiles
Barry Wang, Avi Schwarzschild, Alexander Robey +4
Retrofitting large language models (LLMs) with new behaviors typically requires full finetuning or distillation-costly steps that must be repeated for every architecture. In this w…
Idiosyncrasies in Large Language Models
Mingjie Sun, Yida Yin, Zhiqiu Xu +2
In this work, we unveil and study idiosyncrasies in Large Language Models (LLMs) -- unique patterns in their outputs that can be used to distinguish the models. To do so, we consid…
Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
Mingyu Jin, Kai Mei, Wujiang Xu +5
Large language models (LLMs) have achieved remarkable success in contextual knowledge understanding. In this paper, we show that these concentrated massive values consistently emer…
ConSense: Continually Sensing Human Activity with WiFi via Growing and Picking
Rong Li, Tao Deng, Siwei Feng +2
WiFi-based human activity recognition (HAR) holds significant application potential across various fields. To handle dynamic environments where new activities are continuously intr…
Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
Eungyeup Kim, Mingjie Sun, Christina Baek +2
Recently, Miller et al. (2021) and Baek et al. (2022) empirically demonstrated strong linear correlations between in-distribution (ID) versus out-of-distribution (OOD) accuracy and…