From the 1 of 10 linked papers with an AI index.
10 papers
Sparse Inter-Layer Dependencies of Transformer FFN Neurons
Johannes Knittel, Hanspeter Pfister
The paper introduces a training‑free method to attribute the activation of individual feed‑forward network neurons in Transformers to a small set of upstream neuron activations and…
Bias at the End of the Score
Salma Abdel Magid, Grace Guo, Esin Tureci +4
Reward models (RMs) are inherently non-neutral value functions designed and trained to encode specific objectives, such as human preferences or text-image alignment. RMs have becom…
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
Narmeen Oozeer, Luke Marks, Shreyans Jain +2
Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of…
Task-Specific Directions: Definition, Exploration, and Utilization in Parameter Efficient Fine-Tuning
Chongjie Si, Zhiyi Shi, Shifan Zhang +3
Large language models demonstrate impressive performance on downstream tasks, yet they require extensive resource consumption when fully fine-tuning all parameters. To mitigate thi…
A Rigorous Behavior Assessment of CNNs Using a Data-Domain Sampling Regime
Shuning Jiang, Wei-Lun Chao, Daniel Haehn +2
We present a data-domain sampling regime for quantifying CNNs' graphic perception behaviors. This regime lets us evaluate CNNs' ratio estimation ability in bar charts from three pe…
Can Machines Imitate Humans? Integrative Turing-like tests for Language and Vision Demonstrate a Narrowing Gap
Mengmi Zhang, Elisa Pavarino, Xiao Liu +20
As AI becomes increasingly embedded in daily life, ascertaining whether an agent is human is critical. We systematically benchmark AI's ability to imitate humans in three language…