Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Spectral Guardrails for Agents in the Wild: Detecting Tool Use Hallucinations via Attention Topology
Valentin Noël
Deploying autonomous agents in the wild requires reliable safeguards against tool use failures. We propose a training free guardrail based on spectral analysis of attention topolog…
cs.LG2026
Spectral Archaeology: The Causal Topology of Model Evolution
Valentin Noël
Behavioral benchmarks tell us \textit{what} a model does, but not \textit{how}. We introduce a training-free mechanistic probe using attention-graph spectra. Treating each layer as…
cs.LG2025
Catching Contamination Before Generation: Spectral Kill Switches for Agents
Valentin Noël
Agentic language models compose multi step reasoning chains, yet intermediate steps can be corrupted by inconsistent context, retrieval errors, or adversarial inputs, which makes p…