Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution
Vedant Palit, Florent Draye, Terry Jingchen Zhang +2
Transcoder attribution graphs are usually trained to explain why a model assigns high probability to a particular next token. We introduce Concept-Targeted Attribution (CTA), which…
cs.CL2023
Towards Vision-Language Mechanistic Interpretability: A Causal Tracing Tool for BLIP
Vedant Palit, Rohan Pandey, Aryaman Arora +1
Mechanistic interpretability seeks to understand the neural mechanisms that enable specific behaviors in Large Language Models (LLMs) by leveraging causality-based methods. While t…
cs.CL2023
Knowledge Graph Guided Semantic Evaluation of Language Models For User Trust
Kaushik Roy, Tarun Garg, Vedant Palit +3
A fundamental question in natural language processing is - what kind of language structure and semantics is the language model capturing? Graph formats such as knowledge graphs are…