9 papers
Data-Driven Energy-Based Learning via Gibbs Measures on Hierarchical Structures
L. U. Abdullaev, F. Herrera, U. A. Rozikov +1
We introduce a data-driven probabilistic framework for learning systems based on Gibbs measures on hierarchical structures. Unlike standard empirical risk minimization, where a dat…
High-Dimensional Random Projection for Activation Steering in Language Models
Minh-Hieu Pham, Bach Do, Laziz Abdullaev +2
Activation steering has emerged as a key methodology for controlling the behavior of large language models (LLMs). Existing difference-in-means based methods, however, are fundamen…
Concept Heterogeneity-aware Representation Steering
Laziz U. Abdullaev, Noelle Y. L. Wong, Ryan T. Z. Lee +3
Representation steering offers a lightweight mechanism for controlling the behavior of large language models (LLMs) by intervening on internal activations at inference time. Most e…
Tight Clusters Make Specialized Experts
Stefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev +1
Sparse Mixture-of-Experts (MoE) architectures have emerged as a promising approach to decoupling model capacity from computational cost. At the core of the MoE model is the router,…
Revisiting Transformers with Insights from Image Filtering and Boosting
Laziz U. Abdullaev, Maksim Tkachenko, Tan M. Nguyen
The self-attention mechanism, a cornerstone of Transformer-based state-of-the-art deep learning architectures, is largely heuristic-driven and fundamentally challenging to interpre…
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
Minh-Khoi Nguyen-Nhat, Rachel S. Y. Teo, Laziz Abdullaev +3
Sparse Mixture of Experts (SMoE) has emerged as a promising solution to achieving unparalleled scalability in deep learning by decoupling model parameter count from computational c…