Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
Florent Draye, Abir Harrasse, Vedant Palit +8
Mechanistic interpretability seeks to understand how Large Language Models (LLMs) represent and process information. Recent approaches based on dictionary learning and transcoders…
cs.LG2025
Adaptive Federated Learning Defences via Trust-Aware Deep Q-Networks
Vedant Palit
Federated learning is vulnerable to poisoning and backdoor attacks under partial observability. We formulate defence as a partially observable sequential decision problem and intro…