Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal
Ege Ãakar, Hannah Guan, Kayden Kehe
Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directi…
cs.LG2026
LASER: Low-Rank Activation SVD for Efficient Recursion
Ege Ãakar, Ketan Ali Raghu, Lia Zheng
Recursive architectures such as Tiny Recursive Models (TRMs) perform implicit reasoning through iterative latent computation, yet the geometric structure of these reasoning traject…
cs.LG2025
The Argument is the Explanation: Structured Argumentation for Trust in Agents
Ege Cakar, Per Ola Kristensson
Humans are black boxes -- we cannot observe their neural processes, yet society functions by evaluating verifiable arguments. AI explainability should follow this principle: stakeh…