4 papers
Operationalising the Superficial Alignment Hypothesis via Task Complexity
Tomás Vergara-Browne, Darshan Patil, Ivan Titov +3
The superficial alignment hypothesis (SAH) posits that large language models learn most of their knowledge during pre-training, and that post-training merely surfaces this knowledg…
Tracr-Injection: Distilling Algorithms into Pre-trained Language Models
Tomás Vergara-Browne, Álvaro Soto
Motivated by the surge of large language models, there has been a push to formally characterize the symbolic abilities intrinsic to the transformer architecture. A programming lang…
Eigenpruning: an Interpretability-Inspired PEFT Method
Tomás Vergara-Browne, Álvaro Soto, Akiko Aizawa
We introduce eigenpruning, a method that removes singular values from weight matrices in an LLM to improve its performance in a particular task. This method is inspired by interpre…
Large Language Models are biased to overestimate profoundness
Eugenio Herrera-Berg, Tomás Vergara Browne, Pablo León-Villagrá +2
Recent advancements in natural language processing by large language models (LLMs), such as GPT-4, have been suggested to approach Artificial General Intelligence. And yet, it is s…