1 paper
Akshat Gupta, Jay Yeung, Gopala Anumanchipalli +1
Growing evidence suggests that large language models do not use their depth uniformly, yet we still lack a fine-grained understanding of their layer-wise prediction dynamics. In th…