3 papers
cs.LG2026
Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task
Alicia Curth, Rachel Lawrence, Sushrut Karmalkar +1
We investigate whether transformers use their depth adaptively across tasks of increasing difficulty. Using a controlled multi-hop relational reasoning task based on family stories…
cs.LG2026
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
Victoria Lin, Xinnuo Xu, Rachel Lawrence +4
Despite their strong performance on reasoning benchmarks, large language models (LLMs) have proven brittle when presented with counterfactual questions, suggesting weaknesses in th…
cs.LG2025
Parallel Sampling from Masked Diffusion Models via Conditional Independence Testing
Iskander Azangulov, Teodora Pandeva, Niranjani Prasad +2
Masked diffusion models (MDMs) offer a compelling alternative to autoregressive models (ARMs) for discrete text generation because they enable parallel token sampling, rather than…