Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization
Abdulhady Abas Abdullah, Fatemeh Daneshfar, Seyedali Mirjalili +1
Aligning large language models (LLMs) with human preferences is commonly done via reinforcement learning from human feedback (RLHF) with Proximal Policy Optimization (PPO) or, more…
cs.AI2025
Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
Abdulhady Abas Abdullah, Arkaitz Zubiaga, Seyedali Mirjalili +5
This review surveys the rapid evolution of Meta AI's LLaMA (Large Language Model Meta AI) series - from LLaMA 1 through LLaMA 4 and the specialized parameter-efficient fine-tuning…