collaborators

7 papers

cs.LG2026

When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index

Kleyton da Costa, Bernardo Modenesi

Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic graphs want different attention sh…

cs.AI2026

Perspectives on Tsallis Statistics for Artificial Intelligence

Kleyton da Costa, Bernardo Modenesi

Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter that controls the weight assigned to rare and frequent events. Originally p…

cs.LG2026

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

Seonglae Cho, Zekun Wu, Kleyton Da Costa +3

Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE families remains untested. Single-…

cs.CE2026

GraphNetz: Statistical Benchmarking of Graph Neural Networks with Paired Tests and Rank Aggregation

Kleyton da Costa, Bernardo Modenesi

Graph Neural Networks (GNNs) benchmarks often report single point estimates, even when performance differences are small relative to variation across random seeds, train/test split…

cs.CE2026

Divergence-Guided Particle Swarm Optimization

Kleyton da Costa, Bernardo Modenesi, Ivan F. M. Menezes +1

Particle Swarm Optimization (PSO) is susceptible to premature convergence when the swarm collapses around the global best, particularly on multimodal landscapes in higher dimension…

cs.LG2026

The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

Seonglae Cho, Zekun Wu, Kleyton Da Costa +1

When a language model asserts that "the capital of Australia is Sydney," does it know this is wrong? Models assert misconceptions with the same fluency as facts, so the question ca…