Abstract
n a world where the data is a central piece, we provide a novel technique to design training plans for road cyclists. This study exposes an in-depth review of a virtual coach based on state-of-the- art arti cial intelligence techniques to schedule road cycling training sessions. Together with a dozen of road cycling participants' training data, we were able to create and verify an e-coach dedicated to any level of road cyclists. The system can provide near-human coaching advice on the training of cycling athletes based on their past capabilities. In this case study, we extend the tests of our empirical research project and analyze the results provided by experts. Results of the conducted experiments show that the computational intelligence of our system can compete with human coaches at training plani cation. In this case study, we evaluate the system we previously developed and provide new insights and paths of amelioration for systems based on arti cial intelligence for athletes. We observe that our system performs equal or better than the control training plans in 14 and 24 week training periods where it was evaluated as better in 4 of our 5 test components. We also
we evaluate the system we previously developed and provide new insights and paths of amelioration for systems based on arti cial intelligence for athletes. We observe that our system performs equal or better than the control training plans in 14 and 24 week training periods where it was evaluated as better in 4 of our 5 test components. We also report a higher statistical difference in the results of the experts' evaluations between the control and virtual coach training plan (24 weeks; training load: X 2 = 4.751; resting time quantity: X 2 = 3.040; resting time distance: X 2 = 2.550; ef ciency: X 2 = 2.142). Keywords:virtual coaching; e-coach; arti cial intelligence; training plan; road cycling 1. Introduction Performance is tracked and optimized everywhere at any time. The sports world is an area where one can observe this ever growing need to perform better assessed by the growing market of mobile performance applications (see Deloitte 2017's report). In the business eld, companies try to maximize their bene ts by increasing their performance in multiple levels [1]. The common aspect between these two elds is that one can track the performance and get an idea of it through quantitative information. Data is nowadays ubiquitous, and devices are always collecting data. This makes it easy to access micro- quantitative and visual feedback of our performance but does not necessarily explain the meaning in a macro perspective. Indeed, amateurs as well as professionals in the sports eld are relying on wearable devices as a motivational object, but more importantly to perform better [2]. The growing market of wearables devices enables the tracking of new body-related information to be more precise and more detailed [3]. These ameliorations enable one to get a deeper quantitative feedback, and might also overwhelm the athletes with numerical and complex data, reducing their capabilities to understand and use it properly [4]. This high amount of data can lead the athletes to confusion or misunderstanding and could even make them lose motivation towards using such tools [5]. Making sense of one's data may become a complicated task. The
deeper quantitative feedback, and might also overwhelm the athletes with numerical and complex data, reducing their capabilities to understand and use it properly [4]. This high amount of data can lead the athletes to confusion or misunderstanding and could even make them lose motivation towards using such tools [5]. Making sense of one's data may become a complicated task. The athlete's level will also determine how one interacts and uses the collected data. However, one can fall into overtraining syndrome, which is experienced when a person trains disproportionally Appl. Sci.2021,11, 313.
Appl. Sci.2021,11, 313 2 of 17 and ends up training in an inappropriate way compared to their capabilities and what their body can manage [6]. Additionally, such syndrome can occur when the equilibrium between the training load and the resting time is not respected and can lead to critical situations, both related to physical and mental health [7]. Amateur athletes are also more prone to develop such syndrome, as they do not bene t from coach feedback on the way they train [6]. Personalized training feedback may be seen as a way to mitigate the issues induced by performance tracking in sports. Professional athletes rely on human coaches with expertise in managing training, efforts, and resting time. We suppose that amateur and semi-professional athletes are not interested in the price of a human expert to manage their training as this would be too much for their needs. This level of athletes use training software that lacks a proper personalization component and may experience a decrease in their motivation past a certain time [8]. In our study, we present a novel approach to training management. We leverage on new arti cial intelligence-based work in order to create a virtual coach with training per- sonalization capabilities. The technique we use is able to tailor and take into consideration any level of athletes, as it bases the feedback instructions on the capabilities of its users. Using state-of-the-art technologies and speci c measurements, we ensure a proper and convenient training plan designed and managed around the athletes. Thus, we provide amateurs and semi-professionals with a convenient training coaching system, enabling them to get expert-like instructions and feedback. Additionally, this solution can provide human coaches with a different point of view on their athletes' training and introduce more diversity in training plans. 2. Related Work Through our research work, we reviewed an interesting collection of research papers that demonstrated the trends in computer intelligence applied to sports. We particularly looked at the presence of machine learning and algorithms applied to cycling. The results of the research still yielded additional sports that we comprised in this
and introduce more diversity in training plans. 2. Related Work Through our research work, we reviewed an interesting collection of research papers that demonstrated the trends in computer intelligence applied to sports. We particularly looked at the presence of machine learning and algorithms applied to cycling. The results of the research still yielded additional sports that we comprised in this review. We coded the results of our state-of-the-art review according to Lv. L. and Ye. C.'s categories [9]. We augmented the category set by adding new groups of systems, as reported in Table. Table 1.Applications of machine learning in sports.Applications Number of Papers Training optimization 10 Quantifying athlete's performance 6 Predicting results of athlete 6 Training motivation 4 Injury prevention 2 Rule control 1 Gear settings optimization 1 As a general statement, we report a higher involvement of researchers into systems dedicated to endurance sports, like running [1013] and cycling [1318,1820]. One of the main reasons for this is that these physical activities are quite easy to practice and affordable. Thanks to this ease of access and higher popularity, they both bene t a wider range of sensors that can track the athletes [21]. These sports are also ones that may require less data manipulation in order to work with machine learning. Indeed, rankings in running and cycling competitions are based on incremental scores [14], and one can observe changes in the athletes' performance through the collected timeseries data coming from many different sensors [21]. 2.1. Usable Features for Sports A primary data source for performance measurement is sensors. They can be in- tegrated in speci c devices or by using the ones embedded in smartphones. Raw data
Appl. Sci.2021,11, 313 3 of 17 from these sensors may be used, but they are usually processed in order to remove noise. Sensors provide data with a high frequency and thus may not be used directly. For exam- ple, information such as the 3-axis acceleration enables the understanding of a movement relative to its past state but can highly uctuate across short periods of time since the sensors are particularly sensitive. Most of the time, we observed a higher usage of feature combinations, where features get merged in order to create a new one, a process commonly known as feature extraction. Such measurements are time-related and are highly impacted by past values. In cycling, for instance, one can use Training Stress Score (TSS) to explain the effort required by a training session [22]. Using sensors allows measuring signals without intruding the user's activity and, therefore, without impact on their performance. In contrast, measurements based on self-reporting such as Rating of Perceived Exertion (RPE) require that the athlete provide the required information in an explicit manner. For the RPE, the user should have a certain knowledge of its current capabilities as they have to report the perceived exertion on a 20-point scale. Despite being more accurate than other sensors' measurements, the RPE is hard to use [19]. Systems are no longer dependent on wearable sensors, unlike the ones used for running or cycling. Researchers also calculated statistics based on players or teams' actions and scores, since sports like football where access to a video of matches in high-ranking teams is quite dif cult for external users to access. Thus, statistical data from the matches are used in order to predict future scores [23]. Similarly, in the case of tness, the data gathered on the athlete cannot accurately describe their effort; rather, one needs to use additional devices to track the movements and state on the machine [24,25]. Features explaining the performance of a team or an individual may not be directly related to some sensors data. In fact, one can understand the performance of a football team by checking their scores
gathered on the athlete cannot accurately describe their effort; rather, one needs to use additional devices to track the movements and state on the machine [24,25]. Features explaining the performance of a team or an individual may not be directly related to some sensors data. In fact, one can understand the performance of a football team by checking their scores and enrich the statistics with additional information on whether it was a home win or an away win [26]. External information feeds can also help to understand one's performance, since people share a lot of information through social media platforms. Thus, mining the information shared on social media and using Natural Language Processing (NLP) can provide a rich source of feedback data about an athlete's performance [11]. Sports where the body position has a high impact on performance is also being helped by computer vision capabilities as well as recent research in deep learning. Indeed, it is possible to extract a great number of features and information from a camera feed. The biomechanical data can be treated in order to enhance the movements or correct the postures of the athletes and thus help them perform better. Sports such as golf, tennis or javelin are bene tting from these techniques to track the athlete's position and treat it through image processing [2729]. One's body shape may partially de ne their ability to perform at a certain level, as well as the extent to which one can perform a gesture. We observed the usage of data coming from speci c sensors or measurements made to explain the current organs and physiological status of an athlete. We obviously nd a high usage of the heartrate, despite being criticized for its high variance and dependence on the athlete's form [19]. While heartrate explains the evolution of one's heart beats per minutes, other features are not sensor-based and may require experts to perform measurements. For example, kinanthropometric data is a set of information composed of the body size, its shape or even its composition and may be used to evaluate one's potential performance [30]. In
dependence on the athlete's form [19]. While heartrate explains the evolution of one's heart beats per minutes, other features are not sensor-based and may require experts to perform measurements. For example, kinanthropometric data is a set of information composed of the body size, its shape or even its composition and may be used to evaluate one's potential performance [30]. In Table, we summarize the observed features and classify using our own taxonomy. Compared to outer body-related data, the inner body information may vary faster. The two categories are linked but still separated by a thin line, which is the latter's uctuation across a short period of time. Athlete-related data make a clear categorization of the athlete using their age, sex and anthropometric information. Additionally, computer vision (CV) based features can be extracted from video frames processing. These are measurements based on the skeleton's joint position and enable the understanding of one's body structure and position through time.
Appl. Sci.2021,11, 313 4 of 17 Table 2.The categorization of used data in sports dedicated systems. Sensor-Based Computer Vision Installation Speci c Inner Body Ambient Athlete Related Subjective Measurements Performance Data Injury Related External Sources Computed Values Acceleration Video Displacement of reference point HR Humidity Sex Feelings Phase of play Injury type Social Media TSS Gyroscope X-Factor movement Cable force Strength Location Age RPE Stage of the season Injured side TRIMP Velocity Skeleton joint positions Flexibility Temperature Anthropometric measures Home win Injured body part VO2max Power Technical ef ciency Away win Reoccurrence Exercise economy Distance Force development Results history Days unavailable Anaerobic capability Speed Kinanthropometric measures Average scores pro/con # Injured players Lactate threshold Body mass Overall performance Player availability Fatigue Capillarization Rankings Max strength Oxidative enzyme activity Max speed Endurance performance Diffusion distance
Appl. Sci.2021,11, 313 5 of 17 A high number of sports rely on the athlete's position in time, thus the information may be key to an enhancement in performance [27]. Processing video feeds can provide a high amount of information in professional-level sports, since the data is undisclosed due to the competition aspect. 2.2. Machine Learning Usage in a Sports Context Not all machine learning techniques can be used to treat sports problems. Through the papers we reviewed, we found some trends in the selected techniques. The type of data, the number of features and the distribution of the data are some of the most determinant aspects to consider while choosing a machine learning model. Thus, depending on the features and the quantity of available data, researchers privileged some models among others. In Table, we present the count of machine learning techniques used in the reviewed papers. Table 3.Comparison of machine learning models usage in sports.Techniques Number of Papers ANN 7 SVM 3 CV 3 Clustering 3 Fuzzy logic 3 Naive bayes 3 Feature selection/extraction 2 Ontology 2 K-NN 2 Random forest 1 LogiBoost 1 Arti cial Neural Network (ANN) are the most used approaches of our review. This is due to recent advances in deep learning and ANN-based models, but mainly its gen- eralization capability. ANNs tend to be, when well used, models that can treat any type of data. The architecture is exible, as one can parametrize the number of neurons, the layer types (which depends on the data type), and the starting weights. The only issue is that ANN are known to be overused and, in some cases, may also over t the dataset quite easily. The Support Vector Machine (SVM) approach is very similar to ANN in terms of generalization capabilities. As for ANNs, these models can also support high amounts of data. We are surprised that there are not more papers using this approach. Indeed, the use of SVM ensures avoiding over tting on the available data. On the contrary to supervised machine learning techniques, unsupervised machine learning tries to solve one of the main
in terms of generalization capabilities. As for ANNs, these models can also support high amounts of data. We are surprised that there are not more papers using this approach. Indeed, the use of SVM ensures avoiding over tting on the available data. On the contrary to supervised machine learning techniques, unsupervised machine learning tries to solve one of the main issues of creating a dataset, which is data labelling. The goal of adopting such techniques is generating clusters of data that have similar characteristics or values. The nality of such an approach is mostly an explanation of the dataset's content where one can nd and extract trends or patterns. They seem to be quite ef cient at determining winning strategies in multi-staged competitions [14,15]. We also accounted papers using some older techniques to interpret sports' data, such as fuzzy logic and ontologies. Both techniques provide results that can be easily understood and may also be used with unlabeled data. Fuzzy logic coupled with fuzzy inference can provide an idea of the different states, or classes, of the given data [25]. Thus, one could use it to extract classes from a given dataset that may not be initially understandable. In counterparts, ontologies are used to have a de ned number of states and transitions and one uses semantic reasoning to apply the data to it and extract results. The latter technique will adapt the states and their transition to a dataset, thus it can be used to provide recommendations of another one's performance with a certain degree of exibility [11,17]. Computer vision (CV) is also not directly linked to machine learning techniques and can beused when one needs to extract data from image processing in order to construct a
Appl. Sci.2021,11, 313 6 of 17 statistical analysis or a usable dataset for machine learning techniques [31]. We found many applications (mainly in sports) where it was not possible to use on-bodied sensors to track the user's movements. Thus, the tracking of data such as the skeleton joints can provide information on gestures and postures performed by the athletes and react to them [27]. CV can also be used to track objects' movements and explain speci c behaviors [28]. 2.3. Sports Coaching Based on Machine Learning Virtual coaches can be de ned as, computer systems capable of sensing relevant context, determining user intent and providing useful feedback with the aim of improving some aspect of the user's life [32]. We observed that virtual coaching has evolved as fast as machine learning research, enabling management of larger quantities of information, more data types and offering new models to rely on. However, we hereafter point out aspects linked to coaching and a common gap of all reviewed solutions. E-coaching consists of virtual support for human real-life activities and it can be deployed in a plethora of different contexts, including sports. A coaching support can be provided in many ways, but it is important that the medium chosen for the interaction with the user still needs to be adequate to its application, especially in sports since athletes may not get the information in all contexts (i.e., before, during or after their effort). We relied on the modalities proposed by [33] to construct our synthesis: audio communication; video communication; synchronous text-based communication; asynchronous text-based communication. Audio communication is particularly interesting in sports, since it allows conveying information of the current performance to an athlete without engaging them in a high workload activity. We de ne high workload as any activity where the person needs to think and focus in order to get information. On the contrary, a video communication system represents a high workload during the effort, since the athlete needs to focus on the screen and not on the effort they are making. However, head-mounted displays may reduce the induced workload as the
Description
A novel technique to design training plans for road cyclists using AI.