7 papers
When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index
Kleyton da Costa, Bernardo Modenesi
Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic graphs want different attention sh…
Perspectives on Tsallis Statistics for Artificial Intelligence
Kleyton da Costa, Bernardo Modenesi
Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter that controls the weight assigned to rare and frequent events. Originally p…
Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects
Seonglae Cho, Zekun Wu, Kleyton Da Costa +3
Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE families remains untested. Single-…
GraphNetz: Statistical Benchmarking of Graph Neural Networks with Paired Tests and Rank Aggregation
Kleyton da Costa, Bernardo Modenesi
Graph Neural Networks (GNNs) benchmarks often report single point estimates, even when performance differences are small relative to variation across random seeds, train/test split…
Divergence-Guided Particle Swarm Optimization
Kleyton da Costa, Bernardo Modenesi, Ivan F. M. Menezes +1
Particle Swarm Optimization (PSO) is susceptible to premature convergence when the swarm collapses around the global best, particularly on multimodal landscapes in higher dimension…
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
Seonglae Cho, Zekun Wu, Kleyton Da Costa +1
When a language model asserts that "the capital of Australia is Sydney," does it know this is wrong? Models assert misconceptions with the same fluency as facts, so the question ca…