Showing 2024 · cs.CLShow all
3 papers · 2 filters
cs.CL2024
Interpreting token compositionality in LLMs: A robustness analysis
Nura Aljaafari, Danilo S. Carvalho, André Freitas
Understanding the internal mechanisms of large language models (LLMs) is integral to enhancing their reliability, interpretability, and inference processes. We present Constituent-…
cs.CL2024
Inductive Learning of Logical Theories with LLMs: An Expressivity-Graded Analysis
João Pedro Gandarela, Danilo S. Carvalho, André Freitas
This work presents a novel systematic methodology to analyse the capabilities and limitations of Large Language Models (LLMs) with feedback from a formal inference engine, on logic…
cs.CL2024
Bridging Linguistic Structure and Mechanistic Interpretability for Conceptual Interpretation in Language Models
Nura Aljaafari, Danilo S. Carvalho, André Freitas
Understanding how language models compose meaning from linguistic input remains a central problem in interpretability research. Mechanistic studies have attributed functional roles…