SherpaAI: A Multi-modal Solution for Delivering Personalized and Adaptive Fitness Interventions
arXiv:2604.00968
Abstract
Personalization of exercise routines is a crucial factor in helping people achieve their fitness goals. Despite this, many contemporary solutions fail to offer real-time, adaptive feedback tailored to an individual's physiological states. Contemporary solutions often rely only on static, pre-set plans and rarely adjust in real time to factors such as a user's pain thresholds, fatigue levels, or form during a workout. This work introduces SherpaAI, a multi-modal system that unifies computer vision, physiological sensing (heart rate and voice), and the reasoning capabilities of Large Language Models (LLMs)---modalities that prior systems have largely explored in isolation---to deliver real-time and individually-adaptive guidance across a set of strength, balance, and flexibility exercises. SherpaAI continuously monitors a user's physical form and level of exertion, among other parameters, to provide dynamic interventions focused on exercise intensity, rest periods, and motivation. To validate our system, we performed a technical evaluation confirming our models' accuracy and quantifying pipeline latency, alongside an expert review where certified trainers validated the correctness of the LLM's interventions. Furthermore, in a controlled within-subject study with 25 participants, SherpaAI demonstrated significant improvements over a non-adaptive baseline modeled on typical self-guided workouts (e.g., following online videos or general-purpose AI chat tools for guidance). With SherpaAI, users reported significantly greater enjoyment, a stronger sense of achievement, and significantly lower levels of boredom and frustration. These results indicate that by integrating multi-modal sensing with LLM-driven reasoning, adaptive systems like SherpaAI can create a more engaging and emotionally satisfying workout experience.