activity
20242026
collaborators

6 papers

cs.LG2026

Stronger Normalization-Free Transformers

Mingzhi Chen, Taiming Lu, Jiachen Zhu +2

Although normalization layers have long been viewed as indispensable components of deep learning architectures, the recent introduction of Dynamic Tanh (DyT) has demonstrated that…

cs.LG2025

Command-V: Pasting LLM Behaviors via Activation Profiles

Barry Wang, Avi Schwarzschild, Alexander Robey +4

Retrofitting large language models (LLMs) with new behaviors typically requires full finetuning or distillation-costly steps that must be repeated for every architecture. In this w…

cs.CL2025

Idiosyncrasies in Large Language Models

Mingjie Sun, Yida Yin, Zhiqiu Xu +2

In this work, we unveil and study idiosyncrasies in Large Language Models (LLMs) -- unique patterns in their outputs that can be used to distinguish the models. To do so, we consid…

cs.CL2025

Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding

Mingyu Jin, Kai Mei, Wujiang Xu +5

Large language models (LLMs) have achieved remarkable success in contextual knowledge understanding. In this paper, we show that these concentrated massive values consistently emer…

eess.SP2025

ConSense: Continually Sensing Human Activity with WiFi via Growing and Picking

Rong Li, Tao Deng, Siwei Feng +2

WiFi-based human activity recognition (HAR) holds significant application potential across various fields. To handle dynamic environments where new activities are continuously intr…

cs.LG2024

Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line

Eungyeup Kim, Mingjie Sun, Christina Baek +2

Recently, Miller et al. (2021) and Baek et al. (2022) empirically demonstrated strong linear correlations between in-distribution (ID) versus out-of-distribution (OOD) accuracy and…