2 papers
cs.LG2025
Hyper-SET: Designing Transformers via Hyperspherical Energy Minimization
Yunzhe Hu, Difan Zou, Dong Xu
Transformer-based models have achieved remarkable success, but their core components, Transformer layers, are largely heuristics-driven and engineered from the bottom up, calling f…
cs.LG2024
An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
Yunzhe Hu, Difan Zou, Dong Xu
Deep neural networks have long been criticized for being black-box. To unveil the inner workings of modern neural architectures, a recent work \cite{yu2024white} proposed an inform…