collaborators

5 papers

cs.CL2026

LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference

Hossein Rajabzadeh, Maryam Dialameh, Chul B. Park +2

Autoregressive large language models (LLMs) are bottlenecked by sequential decoding, where each new token typically requires executing all transformer layers. Existing dynamic-dept…

cs.LG2025

Bayesian Mixture of Experts For Large Language Models

Maryam Dialameh, Hossein Rajabzadeh, Weiwei Zhang +2

We present Bayesian Mixture of Experts (Bayesian-MoE), a post-hoc uncertainty estimation framework for fine-tuned large language models (LLMs) based on Mixture-of-Experts architect…

cs.CV2025

EMA-SAM: Exponential Moving-average for SAM-based PTMC Segmentation

Maryam Dialameh, Hossein Rajabzadeh, Jung Suk Sim +1

Papillary thyroid microcarcinoma (PTMC) is increasingly managed with radio-frequency ablation (RFA), yet accurate lesion segmentation in ultrasound videos remains difficult due to…

eess.IV2025

DualSwinUnet++: An Enhanced Swin-Unet Architecture With Dual Decoders For PTMC Segmentation

Maryam Dialameh, Hossein Rajabzadeh, Moslem Sadeghi-Goughari +2

Precise segmentation of papillary thyroid microcarcinoma (PTMC) during ultrasound-guided radiofrequency ablation (RFA) is critical for effective treatment but remains challenging d…

cs.LG2025

ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training

Maryam Dialameh, Rezaul Karim, Hossein Rajabzadeh +5

This paper introduces ECHO-LLaMA, an efficient LLaMA architecture designed to improve both the training speed and inference throughput of LLaMA architectures while maintaining its…