1 paper
Moiz Sadiq Awan, Muhammad Haris Noor, Muhammad Salman Munaf
Automated benchmarks dominate the evaluation of large language models, yet no systematic study has compared user satisfaction, adoption motivations, and frustrations across competi…