3 papers
cs.CL2026
Geometry-Preserving Aggregation for Mixture-of-Experts Embedding Models
Sajjad Kachuee, Mohammad Sharifkhani
Mixture-of-Experts (MoE) embedding models combine expert outputs using weighted linear summation, implicitly assuming a linear subspace structure in the embedding space. This assum…
cs.CL2025
Efficient Large Language Models with Zero-Shot Adjustable Acceleration
Sajjad Kachuee, Mohammad Sharifkhani
Using Large Language Models (LLMs) in real-world applications presents significant challenges, particularly in balancing computational efficiency with model performance. Optimizing…
cs.CL2024
Latency Adjustable Transformer Encoder for Language Understanding
Sajjad Kachuee, Mohammad Sharifkhani
Adjusting the latency, power, and accuracy of natural language understanding models is a desirable objective of an efficient architecture. This paper proposes an efficient Transfor…