Abstract
dividual marathon optimal pacing sparring the runner to hit the “wall” after 2 h of running remain unclear. In the current study we examined to what extent Deep neural Network contributes to identify the individual optimal pacing training a Variational Auto Encoder (VAE) with a small dataset of nine runners. This last one has been constructed from an original one that contains the values of multiple physiological variables for10 differentrunners during a marathon. We plot the Lyapunov exponent/Time graph on these variables for each runner showing that the marathon wall could be anticipated. The pacing strategy that this innovative technique sheds light on is to predict and delay the moment when the runner empties his reserves and ’hits the wall’ while considering the individual physical capabilities of each athlete. Our data suggest that given that a further increase of marathon runner using a cardio-GPS could benefit of their pacing run for optimizing their performance if AI would be used for learning how to self-pace his marathon race for avoiding hitting the wall. Keywords:marathon; deep neural; Lyapunov; hitting the wall; VAE 1. Introduction Marathon running represents one of the most demanding tests of human endurance and strength, requiring participants to demonstrate resilience and mental fortitude over a distance
their pacing run for optimizing their performance if AI would be used for learning how to self-pace his marathon race for avoiding hitting the wall. Keywords:marathon; deep neural; Lyapunov; hitting the wall; VAE 1. Introduction Marathon running represents one of the most demanding tests of human endurance and strength, requiring participants to demonstrate resilience and mental fortitude over a distance of 42.195 km. Elite marathon runners showcase extraordinary capabilities, with the fastest male athletes averaging speeds of approximately 20 km/h and the fastest female athletes achieving around 19.3 km/h. These remarkable feats are not accomplished through a simple, unwavering pace. Instead, they result from carefully calibrated speed fluctuations designed to balance energy conservation and fatigue management. A key component of these performance strategies is the concept of critical speed. The maximum sustainable aerobic pace an athlete can maintain without rapidly accumulating fatigue. Running at or below this pace allows the body to sustain energy supply and meet the muscular demand for oxygen. Exceeding this velocity, however, forces the body’s metabolism to shift towards anaerobic pathways, leading to accelerated fatigue. Through rigorous training, elite runners develop the capacity to approach or even surpass their critical speed for extended periods. By alternating between slower and faster speeds, they optimize their energy use, delay fatigue, and maximize performance [1,2]. This strategy enables them to sustain energy and delay fatigue, thereby optimizing their potential across the race distance without reaching their maximal oxygen uptake (O2max), which cannot be sustained for a long time [3]. The latter represents the maximum AI2025,6, 130 https://doi.org/10.3390/ai6060130
AI2025,6, 130 2 of 22 rate at which an individual can consume oxygen during intense exercise. Higher levels of O2max are typically associated with enhanced endurance and aerobic capacity, enabling runners to maintain faster speeds over extended distances [4]. The combination of an athlete’s O2max and their tolerance to oxygen deficit (their ability to continue exercising when the demand for oxygen surpasses the supply) helps to inform their pacing strategy and predict their endurance potential. In contrast to the elite runners, recreational runners, an increasingly large and diverse demographic often adopt a rigid pacing approach, aiming to maintain a constant speed throughout the marathon. This approach leads them to achieve their O2max [5]. This strategy, however, frequently results in a phenomenon known as the marathon wall, which is characterised by a sudden onset of extreme fatigue that typically occurs around the 26th kilometer [6]. The phenomenon is caused by a depletion of glycogen stores and an increased reliance on less efficient energy sources, which results in a significant reduction in speed. A sharp decline in speed is often observed in recreational runners who encounter this phenomenon, resulting in a final median speed that is considerably lower than their initial pace [7]. This pacing challenge emphasises the necessity for a more adaptive strategy that could assist runners in achieving their optimal performance without experiencing significant energy deficits in the latter stages of the race. As the race progresses, particularly in the final 15 km, physiological indicators such as heart rate, oxygen consumption (O2), and respiratory rate exhibit increasing variability or entropy [8,9]. In this context, entropy reflects the body’s fluctuating physiological state as it strives to meet the escalating demands of prolonged exertion. Two primary types of entropy are relevant to marathon running: 1. Clausius Entropy (Thermodynamic Entropy): Clausius entropy, which is rooted in the principles of thermodynamics, reflects the disorder or heat accumulation within the body during prolonged exertion. As runners approach the final kilometers, their bodies generate a substantial amount of internal heat, particularly in the presence of elevated temperatures, with body temperatures frequently reaching over 40 ◦ C.
marathon running: 1. Clausius Entropy (Thermodynamic Entropy): Clausius entropy, which is rooted in the principles of thermodynamics, reflects the disorder or heat accumulation within the body during prolonged exertion. As runners approach the final kilometers, their bodies generate a substantial amount of internal heat, particularly in the presence of elevated temperatures, with body temperatures frequently reaching over 40 ◦ C. This rise in thermodynamic entropy can place additional strain on the body’s cooling and energy systems, thereby exacerbating fatigue [10,11]. 2. Shannon entropy (informational entropy) is defined as follows: Derived from infor- mation theory, Shannon entropy refers to the amount of predictability in the runner’s physiological data. During the marathon, there is a reduction in Shannon entropy in the final 10 km, which indicates that the physiological responses become less varied and more predictable as the body reaches a state of fatigue. This reduction in infor- mational entropy indicates that the body’s capacity to adapt in a dynamic manner is impaired, thereby reducing the efficacy of self-regulatory mechanisms in the latter stages of the race [9]. One promising avenue for enhancing pacing strategies in real time is the use of a variable automatic encoder. This AI-based technology is capable of analysing intricate physiological signals and translating them into actionable feedback for the runner. By dynamically encoding data from variables such as heart rate, O2, speed, and perceived exertion, a variable automatic encoder can provide a continuously updated representation of the runner’s state. Integration of this encoder with a cardio-GPS device could facilitate the provision of personalized pacing adjustments to runners based on their real-time physical state. This would assist them in optimising speed fluctuations without exceeding their limits, and in developing an understanding of the relationship between their perceived exertion and their actual physiological response in terms of % of O2max. Indeed, AI could assist with pacing in a manner that is commensurate with the Borg Scale for Perceived Exertion [12]. In structured training program, runners are frequently introduced to a variety of levels of exertion, which are typically evaluated using the Borg scale of perceived
and their actual physiological response in terms of % of O2max. Indeed, AI could assist with pacing in a manner that is commensurate with the Borg Scale for Perceived Exertion [12]. In structured training program, runners are frequently introduced to a variety of levels of exertion, which are typically evaluated using the Borg scale of perceived
AI2025,6, 130 3 of 22 exertion. The scale ranges from 6 to 20, with each level corresponding to a subjective feeling of effort and difficulty. The most commonly utilized levels for marathon training are as follows [13]: The pace is designated as “Easy” and is equivalent to a speed of 5.5–6.4 km/h. The individual is comfortable and able to engage in conversation. The 14th level of exertion, designated as “moderate pace”, is characterized by a moderate level of effort that allows for sustained conversation. The level of difficulty is moderate, yet the pace is sustainable. The seventeenth level is designated as “Hard Pace.” This level of exertion is consider- able, yet it can be sustained for relatively brief distances. The Maximal Effort (20): This level of exertion is close to the point of exhaustion and is therefore not sustainable over extended periods of time. Perceived exertion levels allow runners to adjust their pace according to the specific demands of a race. This study aims to determine whether an AI system can learn these exertion levels through calibration tests to establish a personalized energetic signature. Based on O2max, oxygen deficit tolerance, and perceived exertion, this signature could guide AI-assisted pace adjustments during the race. By analyzing sophisticated physiologi- cal data, this pilot study explores whether an AI-powered system can offer more effective pacing strategies for marathon runners than traditional cardio-GPS devices. The goal is to design an adaptable pacing assistant that integrates and encodes each runner’s unique physiological responses, enabling optimal pace adjustments while avoiding the rigidity of fixed speed maintenance. Such an approach could allow recreational runners to input personalized pacing profiles into their devices, with AI guiding them through optimal speed variations across a marathon distance. This innovation aims to support runners in achieving their personal best under safe, optimized pacing conditions, fostering greater enjoyment and sustainability, especially for diverse and aging populations. To demonstrate this potential, a pilot experiment investigates the capacity of a deep neural network to extract an energetic signature. Physiological modulations in a cohort of runners during a marathon were analyzed using extensive datasets, including
to support runners in achieving their personal best under safe, optimized pacing conditions, fostering greater enjoyment and sustainability, especially for diverse and aging populations. To demonstrate this potential, a pilot experiment investigates the capacity of a deep neural network to extract an energetic signature. Physiological modulations in a cohort of runners during a marathon were analyzed using extensive datasets, including heart rate and speed (Garmin 630), oxygen uptake (O2), respiratory frequency, and metabolic data (Cosmed K5). The use of deep learning in multi-sensor sports data analysis remains uncommon, partly due to the high cost and variability of data acquisition. To address this, innovative strategies such as fractal methods and data augmentation techniques, including sliding windows, were applied to enhance temporal progression within datasets. The study focuses on analyzing marathon runners’ physiological parameters through deep neural networks to derive insights about performance and propose race strategy ad- justments. Artificial intelligence can assist sports physiologists in prioritizing performance tests, uncovering details that are otherwise inaccessible to the human eye. A Variational Autoencoder (VAE), a generative statistical model, was used to create individual signatures sensitive to physiological variations and fatigue [14]. Additionally, Hölder exponents and multifractal spectrum analysis provided a deeper understanding of cardiac autoregulation during intense exercise [15]. While Lyapunov exponents have been used to characterize equilibrium plateau [16], their integration into a multivariable energetic context remains unexplored and could help identify exhaustion points and unsustainable pacing with greater precision. The study seeks to elucidate the unique physiological signatures of marathon runners and explore the interpretability of Garmin and K5 data. Ultimately, it aims to detect fatigue-induced disruptions in race dynamics, advancing our understanding of marathon performance and supporting improved training and race strategies.
AI2025,6, 130 4 of 22 2. Materials and Methods 2.1. Subjects Even if we started with 10 runners, one of them was excluded on account of incomplete data. This was due to a malfunction in the analyzers (a battery issue) that occurred at the half-marathon. Then, nine recreational but well-trained male marathon runners (mean age: 40.1±10.6 years ; weight: 72.7±6.5 kg; height: 178.3±7.5 cm) with performance representative of the average performance of non-elite runners but well-trained male marathon runners whose performance are in the first quartile of performance in popular marathon as the Paris marathon [17] (Table analysis, we deliberately included only one gender in our investigation. Table 1.Subjects Age, Personal Best Marathon Time and the Performance at the Sénart Marathon. * Best Personal Time reached at the Sénart Marathon. N ◦ Runners Age (Years) Faster Marathon Time (Years) Sénart Marathon (2019) 1 47 03 h 12 ′ 48 ′′ (2016) 03 h 31 ′ 34 ′′ 2 44 03 h 34 ′ 57 ′′ (2019) 03 h 34 ′ 57 ′′ * 3 22 03 h 22 ′ 40 ′′ (2019) 03 h 22 ′ 40 ′′ * 4 34 02 h 50 ′ 00 ′′ (2019) 02 h 50 ′ 00 ′′ * 5 47 02 h 59 ′ 22 ′′ (2016) 03 h 32 ′ 07 ′′ 6 58 03 h 27 ′ 32 ′′ (2013) 04 h 30 ′ 34 ′′ 7 29 02 h 57 ′ 03 ′′ (2015) 03 h 14 ′ 24 ′′ 8 36 03 h 27 ′ 58 ′′ (2017) 03 h 51 ′ 44 ′′ 9 43 02 h 44 ′ 00 ′′ (2015) 03 h 13 ′ 53 ′′ All participants volunteered and maintained their regular training routines without alter- ations. The selected runners had prior experience completing a minimum of two marathons. They had been engaged in consistent training, involving three to four sessions per week, cover- ing a range of 50 to 80 km per week, for over 5 years. Every week, the participants incorporated a High-Intensity Interval Training session, involving 6 repetitions of 1000 m at
alter- ations. The selected runners had prior experience completing a minimum of two marathons. They had been engaged in consistent training, involving three to four sessions per week, cover- ing a range of 50 to 80 km per week, for over 5 years. Every week, the participants incorporated a High-Intensity Interval Training session, involving 6 repetitions of 1000 m at intensities between 90% to 100% of their maximal heart rate, along with a tempo training session of 15 to 25 km at speeds ranging from 100% to 90% of their average marathon pace. Ethical considerations were met, as the study’s objectives and procedures received approval from an institutional review board (CPP Sud-Est V, Grenoble, France; reference: 2018-A01496-49). All participants were well-informed about the study and provided written consent to participate. Table participants’ ages, their personal best marathon completion times, and the year these performances were achieved. Notably, some of the runners achieved their personal best during the Sénart Marathon. Table 2.Hyperparameters chosen for VAE-K5 (K5 Cosmed) and VAE-GAR (Garmin watch). The model was trained on a GPU-equipped system NVIDIA RTX 3090 with 32 GB RAM, and each training run took approximately 2 h for convergence. The VAE architecture consists of two fully connected encoder and decoder layers, with a latent dimension of 24 or 12, optimized using the Adam optimizer with a learning rate of 10-4. VAE-K5 VAE-GAR input size 60 60 number of samples 10,844 10,844 Latent dim 24 12 Hidden Layers 3 (encoder) 3 (encoder) Batch Size 32 64 Learning rate 10-4 10-4 KL div. weight 10-4 10-3 Epochs 350 200
AI2025,6, 130 5 of 22 2.2. The Marathon and Experimental Measures In the context of the marathon event, all participants participated in an official race known as the Sénart Marathon in France. The race commenced at 9 a.m. under specific environmental conditions. On 1 May 2019, in Sénart, the weather included temperatures between 11 and 15 ◦ C (from 9 a.m. to 1 p.m.), no precipitation, and an average humidity of 60%. Blood lactate levels were assessed using a finger-based lactate measurement device (Lactate PRO2 LT-1730; ArKray, Kyoto, Japan) immediately after a 15-min warm-up at a leisurely pace and three minutes after crossing the finish line. Throughout the study, we collected continuous data on respiratory gases (oxygen up- take [O2], ventilation [E], and respiratory exchange ratio [RER]) using a portable telemetric system (K5; Cosmed, Rome, Italy) that allowed breath-by-breath analysis. Additionally, a combination of a global positioning system (GPS) watch (Garmin, Olathe, KS, USA) and the K5 system was utilized to monitor heart rate and speed responses, with 5-s averaged data, during each trial. The data collection sampling frequency were the following: – Garmin data: Collected at a frequency of 1 Hz (once per second), providing continuous recording of the athlete’s pacing, heart rate, and GPS coordinates. – Cosmed K5 data: Acquired at a sampling rate of 0.2 Hz (once every 5 s), capturing breath-by-breath physiological parameters, including oxygen uptake (VO2), carbon dioxide production (VCO2), and ventilation (VE). – Outlier Treatment: Outliers were identified using statistical criteria based on phys- iological plausibility, particularly for heart rate, VO2, and speed data. Heart rate values outside physiologically plausible ranges (below 40 bpm or above 220 bpm) were identified as outliers. VO2and speed data points were similarly screened using statistical thresholds (values beyond three standard deviations from the participant’s mean). Approximately 1.5% of the collected data points were identified as outliers and subsequently removed from the dataset. – Missing Data: Missing data (2% of the dataset) ocasionally arise due to sensor failures, communication issues with the athlete, or manual data-collection errors. Missing data were addressed using spline interpolation when the gap was short
three standard deviations from the participant’s mean). Approximately 1.5% of the collected data points were identified as outliers and subsequently removed from the dataset. – Missing Data: Missing data (2% of the dataset) ocasionally arise due to sensor failures, communication issues with the athlete, or manual data-collection errors. Missing data were addressed using spline interpolation when the gap was short (less than 30 s), connecting the last known data point to the next known data point using polynomials for a smoother approximation. Longer gaps or substantial missing segments were excluded from the analysis to ensure data integrity. Kalman filter from Subway and Stoffer [18] were used in this paper for ensuring each data point is assigned one definitive imputed value. – Scaling of Data: for any variable (e.g., heart rate), the standardized value calculated as, where denote the mean and standard deviation of x over all observations in the dataset, All physiological variables–heart rate, VO2, and speed–were standardized using Z-score normalization before being provided as inputs into the VAE. Such uniform scaling also assists in dimensionality reduction and network training by preventing any one variable’s scale from dominating the learning process. To prevent pacing-related influences, runners were encouraged to self-pace their runs while the cardio-GPS display was concealed. Hydration and refreshment points were available every 5 km during the marathon, with additional stations offering sponges and sustenance every 5 km from 7.5 km onward. Runners could remove their masks at these points to consume food and beverages. Con- sistently, all participants drank one glass of water and consumed fruit at each hydration station along the route and at the Start/Finish area. For metabolic assessment during exercise, the participants utilized COSMED reusable face masks constructed from silicone to prevent allergenic reactions. These masks were ergonomically designed to fit snugly and comfortably, maintaining a proper seal without
AI2025,6, 130 6 of 22 compromising data accuracy. In high-intensity exercises, the runners used masks with inspiratory valves to reduce resistance during inhalation, enhancing comfort. The rate of Perception Exertion (RPE) was tracked using the Borg 6–20 scale [12], and participants reported their level of fatigue at least every kilometer or more frequently as needed. This scale was employed to correlate physiological stress indicators with marathon fatigue, and participants were familiarized with it during the two weeks preceding the race. 2.3. Mathematical Procedure 2.3.1. Variational Auto Encoder Autoencoders are a class of neural networks widely employed in unsupervised learn- ing tasks, particularly in the domain of dimensionality reduction, feature learning, and data generation. The fundamental architecture of an autoencoder consists of an encoder and a decoder, which work in tandem to learn a compressed representation of the input data. The encoder maps the high-dimensional input data into a lower-dimensional latent space, capturing essential features and patterns. Subsequently, the decoder reconstructs the original input data from this compressed representation (Figure). (a) (b) Figure 1.Principle of Autoencoders (a) Illustration of Variational Autoencoder (VAE) model ar- chitecture. (b) Illustration of Autoencoder (AE) model architecture; (b) Illustration of Variational Autoencoder (VAE) model architecture. The encoder and decoder components of a standard autoencoder are typically im- plemented using feedforward neural networks. During training, the autoencoder aims to minimize the difference between the original input and the reconstructed output. This pro- cess encourages the network to learn a meaningful encoding of the input data in the latent space. Autoencoders have found applications in various fields, including image denoising, anomaly detection, and feature extraction. While conventional AEs are effective at learning compact representations of data, they lack a probabilistic interpretation, making them limited in tasks that require uncertainty estimation or generative capabilities. Variation Autoencoders (VAEs) address this limitation by introducing probabilistic modeling into the autoencoder framework. VAEs reinterpret the latent space as a probability distribution, allowing for the generation of new data samples by sampling from this distribution. The key principle of VAEs is to impose a constraint on the latent space’s distribution to encour- age it
Description
This study examines how deep neural networks can optimize pacing strategies for recreational marathon runners.