1 citations · 1 across the 15 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
Roy Miles, Aysim Toker, Andreea-Maria Oncescu +3
Reasoning with large language models often benefits from generating multiple chains-of-thought, but existing aggregation strategies are typically trajectory-level (e.g., selecting…
cs.CL2024
From Attention to Activation: Unravelling the Enigmas of Large Language Models
Prannay Kaul, Chengcheng Ma, Ismail Elezi +1
We study two strange phenomena in auto-regressive Transformers: (1) the dominance of the first token in attention heads; (2) the occurrence of large outlier activations in the hidd…