2 papers
cs.CR2026
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
Matteo Leonesi, Francesco Belardinelli, Flavio Corradini +1
Alignment faking (AF) occurs when an LLM strategically complies with training objectives to avoid value modification, reverting to prior preferences once monitoring is lifted. Curr…
cs.FL2024
The Language for Programming Graph Neural Networks
Matteo Belenchia, Flavio Corradini, Michela Quadrini +1
Graph neural networks form a class of deep learning architectures specifically designed to work with graph-structured data. As such, they share the inherent limitations and problem…