4 papers
Optimizing Against Safety Representations: Activation-Guided Adversarial Suffixes and the Geometry of Refusal
Ege Ãakar, Hannah Guan, Kayden Kehe
Behavioral alignment in large language models often masks fragile internal safety representations. Recent work suggests that refusal behavior is mediated by low-dimensional directi…
LASER: Low-Rank Activation SVD for Efficient Recursion
Ege Ãakar, Ketan Ali Raghu, Lia Zheng
Recursive architectures such as Tiny Recursive Models (TRMs) perform implicit reasoning through iterative latent computation, yet the geometric structure of these reasoning traject…
Boule or Baguette? A Study on Task Topology, Length Generalization, and the Benefit of Reasoning Traces
William L. Tong, Ege Cakar, Cengiz Pehlevan
Recent years have witnessed meteoric progress in reasoning models: neural networks that generate intermediate reasoning traces (RTs) before producing a final output. Despite the ra…
The Argument is the Explanation: Structured Argumentation for Trust in Agents
Ege Cakar, Per Ola Kristensson
Humans are black boxes -- we cannot observe their neural processes, yet society functions by evaluating verifiable arguments. AI explainability should follow this principle: stakeh…