3 papers
cs.LG2026
Marginals Before Conditionals
Mihir Sahasrabudhe
We construct a minimal task that isolates conditional learning in neural networks: a surjective map with K-fold ambiguity, resolved by a selector token z, so H(A | B) = log K while…
cs.CL2025
Directional Optimization Asymmetry in Transformers: A Synthetic Stress Test
Mihir Sahasrabudhe
Transformers are theoretically reversal-invariant: their function class does not prefer left-to-right over right-to-left mappings. Yet empirical studies on natural language repeate…
cs.CL2025
Reversal Invariance in Autoregressive Language Models
Mihir Sahasrabudhe
We formalize a structural property of the causal (autoregressive) language modeling (CLM) objective: reversal invariance. Formally, the next-token prediction loss assigns identical…