← Back to library
article 2024 7 pages

RUNNING POWER: METHODS FOR EVALUATING METABOLIC CAPACITY

Ivanka Karparova, Dimitar Dimitrov

Journal
Research in Kinesiology
Publication type
Original paper
Study type
longitudinal case study
Population
recreational runners

Abstract

This longitudinal case study presents the use of Stryd™ Running Power, as a sufficiently precise surrogate of oxygen measurements, thus a valid way to obtain someone’s second threshold via Critical Power (CP) estimation. There are several CP models available. Our goal was to evaluate the similarity between the following CP (or FTP) models - Stryd™ PowerCenter, TrainingPeaks ™ WKO5™, GoldenCheetah , Intervals.icu, and 2-parameter CP calculation. We are using 4 years' worth of Running Power data for a recreational runner, who has many max efforts under 30 minutes to supply the models with accurate data. We also have a few unintended 3-minute all -out (3MT) tests during 5k races. The 3MT is already a validated method to estimate CP and it can give us one more way to review the correctness of the models. We hypothesize that the data collected from the technological devices and based on the tests can be a good enough descriptor of endurance performance. Keywords: Running Power, Critical Power, Anaerobic Work Capacity, Functional Threshold Power INTRODUCTION There is a range of intensities for which exercise can be maintained for a long time (hours) and once we go above this range our time to exhaustion becomes minutes. This transition point has been extensively studied and there are several concepts associated with it, one of them is Critical Power (Poole, Rossiter, Brooks & Gladden, 2021) . A century ago, some of the pioneers of Exercise Physiology, namely prof Archibald V. Hill and his colleague Harley Lupton introduced the terms: “Maximal Oxygen Consumption”, “Steady State” and “Oxygen debt” (Hale, 2008). Hale’s review of Hill’s work goes into details of the evolution that happened over time and nuances that Hi ll and Lupton’s successors elucidated, however, the amount of focus given to Oxygen Debt and Steady State is minuscule, compared to the focus given to Maximal Oxygen

Lupton introduced the terms: “Maximal Oxygen Consumption”, “Steady State” and “Oxygen debt” (Hale, 2008). Hale’s review of Hill’s work goes into details of the evolution that happened over time and nuances that Hi ll and Lupton’s successors elucidated, however, the amount of focus given to Oxygen Debt and Steady State is minuscule, compared to the focus given to Maximal Oxygen Consumption (VO2Max). While the pioneers have not pinpointed a method of identifying the (Maximum) Steady State, this appears to have been done in the work, published by Monod and Scherrer in 1960, originally in French and some years later in English. For a synergic muscle group, they performed multiple dynamic efforts to fatigue in the 2–30-minute range. Using the collected data, they estimated “the threshold of local exhaustion”, producing a formula for work limit: Wlim=a+b*tlim (Scherrer & Monod, 1960). The authors conclude that it is easy to interpret this equation. Everything happens as if the maximum work results from the use of an energetic reserve (a) and energy of reconstitution, the rate of which cannot exceed the maximum (b) (Monod & Scherrer, 1965). Another of their conclusions about CP has not stood the test of time, namely that “when the applied power is less than or equal to the critical power” (according to the previous equation) depletion cannot occur (Poole, Burnley, Vanhatalo, Rossiter & Jones, 2016) . Critical power denotes the greatest rate of energy transduction (oxidative ATP resynthesis) sustainable without continuously depleting Wʹ (energy store component expressed in kJ). Exhaustion (time t) occurs when Wʹ is fully depleted or unavailable (Craig, Vanhatalo, Burnley, Jones & Poole, 2019). Above a certain individual power, intolerance to pulmonary oxygen uptake and maintenance of lactate in a steady state leads to an increase in fatigue and cessation of exercise. For everyone, there is a range of intensities in which one is able or not to maintain the stability of the processes. Another way to determine the threshold intensity is by performing several high- intensity tests. Through these tests and by determining the power output curve, which reaches a plateau within 2-3 min of

in fatigue and cessation of exercise. For everyone, there is a range of intensities in which one is able or not to maintain the stability of the processes. Another way to determine the threshold intensity is by performing several high- intensity tests. Through these tests and by determining the power output curve, which reaches a plateau within 2-3 min of the start of the maximum test, the individual's tolerance to endurance exercises can be determined with great accuracy. For everyone, the power–duration relationship is a hyperbolic function (with an asymptote known to us as the critical power) and a curvature constant called W` (kJ) (or D` in meters in running). The relationship is described by the equation T=w'/(P-CP). The asymptote of the hyperbolic relation between external power and time to task failure, critical power, represents the threshold intensity above which systemic and intramuscular metabolic homeostasis can no longer be maintained. (Goulding & Marwood, 2023). With the increasing of wearable data and tools to analyze this data, is no longer necessary to calculate asymptotes manually this can be performed by algorithms. The goal of this study is to compare those platforms. We used four years' worth of training data for one athlete. A selected number of those were analyzed to establish which Stryd™ metrics are affected the most by changes in Running Effectiveness ( Karparova & Dimitrov, 2023). He also has several 1k (or similar) max efforts, as well as many in between, which lets us review the CP estimates of various models and the amount of "debt" (W') that is estimated from the different models. The “anaerobic” parameter W’ is generally described as a battery, which can be used once we go above the CP intensity and recharge, once we slow down under it (Skiba & Clarke, 2021).

31 The “aerobic” parameter, CP, is an estimation of the Maximal Lactate/Metabolic Steady State (MLSSwork/MLSSvelosity) and represents a metabolic rate, above which you fatigue faster (minutes) and below which you fatigue slower (hours) (Billat, Sirvent, Koralsztein & Mercier, 2003). The Time to Fatigue/Exhaustion (TTF/TTE) at this “quasi-steady state” intensity typically cited is in the 30–60-minute range, which has been demonstrated to be highly trainable via specific protocol in a running cohort like our subject - 10k time around 40 minutes (Billat, Sirvent, Lepretre & Koralsztein, 2004). STRYD AND RUNNING POWER The introduction of Running Power in the mid-2010s allowed athletes and coaches to have similar concepts, vocabulary, and tools in c ycling and running, thus easing the exchange of knowledge and ideas. Cyclists are generally more familiar with the term Functional Threshold Power (FTP), which has several oversimplified descriptions and tests (95% of 20 min power, 60 min power, etc.), but the current complete FTP model has 5 parameters (one of which is mFTP) and is derived from the software TrainingPeaks™ WKO5™, we are only looking into the “aerobic” mFTP and “anaerobic” Functional Reserve Capacity (FRC), as direct successors of CP and W’ respectively. The original study by Monod and Scherrer clearly illustrated that different synergistic muscle groups have different values for parameters A (now known as W’) and B (now known as CP), thus it is erroneous to assume that athletes’ Running CP and Cycling CP would be the same number, though same web platforms (Strava™, TrainingPeaks™ web/mobile, as of March 2024) allow users to set only one FTP/CP number. The WKO5™ software, however, has different values for each sport, and so do intervals.icu and GoldenCheetah. We will not be delving into the differences between Running and Cycling power values of the same athlete in this study, but it is worth mentioning the limitations of the available platforms. To put it simply - the Stryd foot pod is a multi-sensor device, that measures the foot movement in all directions (at can now visualize the full footpath in later revisions); using their proprietary algorithm they estimate the number of Watts

values of the same athlete in this study, but it is worth mentioning the limitations of the available platforms. To put it simply - the Stryd foot pod is a multi-sensor device, that measures the foot movement in all directions (at can now visualize the full footpath in later revisions); using their proprietary algorithm they estimate the number of Watts per kilogram, which are required to move our body at the measured number of meters per second. This value is later multiplied by the weight value and recorded in the file, which is what we pass through the different algorithms. The general suggestion they have is to “set and forget” the Weight value, this was the case for our runner, who in 2019 set the weight to 90 kg, even though his better performances in 2020 came at an actual weight of 87 kg. Weight was measured every morning with the same Garmin™ Index™ scale and auto -uploaded to their web platform. From November 2021 onward the Stryd weight setting was updated every few months before every race or formal test to match the nearest kilogram the morning of. The value excludes extra weight from subsequent food/water and clothes/shoes. From a power estimation point of view: an actual weight lower than the value set in Stryd will inflate the estimated power (we are moving less weight than declared) - this will be visible with the 2020 CP data and power vs heart rate data; while the reverse scenario will deflate the estimated power (we move more weight, compared to what Stryd knows about). This can be a consideration in ultramarathons, where running with a few extra kilograms of water, food, and gear is common. It is also a consideration, if we want to review our progress (or regress) and figure out if the change originates from weight change, (de)training of our internal metabolic abilities, or if it is a change in Running Effectiveness. There is already literature on runners using Stryd™ power data and comparing them with lab-based VT2 (Ventilatory threshold 2) and OBLA (Onset of blood lactate accumulation 4

we want to review our progress (or regress) and figure out if the change originates from weight change, (de)training of our internal metabolic abilities, or if it is a change in Running Effectiveness. There is already literature on runners using Stryd™ power data and comparing them with lab-based VT2 (Ventilatory threshold 2) and OBLA (Onset of blood lactate accumulation 4 mmol), which is outlined. However, individual runner variance between CP STRYD and other CP models must be a consideration for runners and coaches. CP STRYD was most similar in intensity to VT2 and OBLA and was predictive of outdoor and laboratory running performance (Dearing & Paton, 2023). To cover these considerations for runners and coaches we are presenting the current longitudinal comparison of the commercially available models, but for the same runner. The models under test are: ● CP2: CP calculation, based on 2 max efforts: short (3- 4 min) and long (10- 23 min). ● mFTP: from TrainingPeaks™ WKO5™ software. ● GC CP: from the open-source software GoldenCheetah. ● Stryd CP: The output of the Stryd PowerCenter model. ● eFTP: the value from the intervals.icu web platform. Additionally, several 5k races were done fatigued and some of them turned into involuntary (extensive) 3-Min аll-оut Tests (3MT), which is already established as a valid way to estimate Critical Power in both cycling and running (Vanhatalo, Doust & Burnley, 2007; Pettitt, Jamnick & Clark, 2012). We will also present the ratio of Power to Heart Rate, which TrainingPeaks™ has named the Efficiency Factor (EF) of sub-maximal efforts, and illustrate how the EF curve has shifted over time, which is - we get higher power for the same heart rate. METHODS In pursuit of improving performance, testing, and monitoring take part in the daily work that coaches and athletes exert to understand the resulting adaptions to training better (Ruiz- Alias, Ñancupil-Andrade, Pérez-Castilla & García-Pinillos, 2023). There is an existing study, showing how different Critical Power Models compare with multiple subjects and cycling, here the focus is longitudinal, single subject, and the commercially available models from Stryd and WKO5 as well as free

the daily work that coaches and athletes exert to understand the resulting adaptions to training better (Ruiz- Alias, Ñancupil-Andrade, Pérez-Castilla & García-Pinillos, 2023). There is an existing study, showing how different Critical Power Models compare with multiple subjects and cycling, here the focus is longitudinal, single subject, and the commercially available models from Stryd and WKO5 as well as free models, such as GoldenCheetah and intervals. icu (Bergstrom, Housh, Zuniga, Traylor, Lewis, Camic, Schmidt & Johnson, 2014). For our case study, the runner has been doing all his runs since September 2019 using the Stryd™ Wind (Sv3) foot pod, which in December 2022 was upgraded to Stryd™ Next Gen (Sv4). Karparova and Dimitrov (2023) found in their last research that there are minor differences in the metrics reported by both models, which was problematic when trying to review Running Effectiveness but for this multi-model analysis of CP, these differences are irrelevant. The study is not blinded - the test subject has been using manual 2-part tests as well as PowerCenter and WKO5 to be aware of his second threshold estimation (CP/FTP), which makes the models well-maintained for the full 4- year period, whereas more random data could yield a higher amount of difference across the CP models. Since there does not appear to be a study comparing the commercially available models against each other and we already had the data in a few of them, we completed the picture by adding Golden Cheetah and intervals.icu to the mix.

32 RESULTS AND DISCUSSION Table 1. presents the values for the two-parameter model. It should be noted that both short and long power values follow the same general progression regardless of weight. Table 1. : Data from all formal 2-parameter tests, along with the Stryd weight setting and the CP estimations Short: 1k / 3 min Long: 10 min / 5000m CP2 Time (m:s) Power (W) Weight (kg) Time (m:s) Power (W) Weight (kg) CP (w/kg) W' (J/kg) CP (W) W' (kJ) December 2019 03:00 422 90 10:00 363 90 3.75 169 337 15.2 August 2020 03:36 421 90 22:52 342 90 3.64 221 327 19.8 March 2021 03:11 423 80 19:34 353 80 4.24 200 339 15.9 October 2021 03:16 430 80 19:16 363 80 4.37 198 349 15.8 February 2022 03:29 424 83 19:51 364 83 4.23 183 351 15.2 October 2022 03:27 431 85 19:30 378 85 4.31 157 366 13.3 Jan 2023 03:00 438 84 19:04 379 84 4.43 152 368 12.6 May 2023 03:20 438 83 19:41 373 83 4.33 189 360 15.6 September 2023 03:16 438 81 19:23 376 81 4.49 177 363 14.4 January 2024 03:25 420 82 19:33 370 82 4.38 151 359 12.4 We haven't found any information that anyone has done a cross-platform comparison of algorithms. Since this is a brand new thing, there is room for improvement. There is decent connectivity between different platforms, and the process of obtaining so much data that was not previously planned to be used in a study, in all places came with some complications that are worth describing. Most of the running data is collected with Garmin watches (Fenix 5x, Fenix 6x, Forerunner 245, Coros Pace 2) and is available on their platform. Both Garmin Connect and the Coros Training Hub were connected to several web platforms, and upon completion, each training file was automatically uploaded to TrainingPeaks, Stryd, and Strava (and a few others). WKO5 automatically downloads the data from TrainingPeaks. The intervals.icu platform was added much later to the mix, but it allows all data to be downloaded from Strava and

Both Garmin Connect and the Coros Training Hub were connected to several web platforms, and upon completion, each training file was automatically uploaded to TrainingPeaks, Stryd, and Strava (and a few others). WKO5 automatically downloads the data from TrainingPeaks. The intervals.icu platform was added much later to the mix, but it allows all data to be downloaded from Strava and the last 1 year from Garmin, after which the data is combined. For the research, we requested a complete archive from the Garmin Connect database. Imports are processed automatically at intervals.icu. For GoldenCheetah, we used the same archive file, removing some of the invalid files. A small number of files have both Garmin Running Power (values that are about 25-30% higher than Stryd) which can confuse the algorithms, this file channel has been removed in WKO5, intervals.icu and GoldenCheetah; Stryd Powercenter is not affected by this type of data, nor is the manual Two Maximum Effort (CP2) methodology. All adjustments have been made to the platforms for adjusting the athlete's weight. COLLECT CP/FTP VALUES ● CP 2: calculated using the web tool https://superpowercalculator.com/calculators/6, which is a simplified copy of the SuperPowerCalculator in Google Sheets, created by and used by the „Run with Power“ community. We assume that the values have not changed until the next official test. ● StrydCP: their web platform shows the changes in your CP values over time in a chart called My Training (which can be expanded to show many years). Each value was added to the table. Stryd does not provide a value to W'. ● mFTP/FRC: TrainingPeaks has a "PD History Metrics Chart" that has been changed to Report and the mFTP/FRC values for the weeks of interest have been added to the CP2 and Stryd values. ● GoldenCheetah CP/ W': We manually added the CP and W' software values to the central table, with all other metrics. ● intervals.icu: Fitness, created a new view covering the full-time range and presenting a Run FTP chart. We couldn't find a way to copy all the values, so just like with the Stryd Powercenter - the

and Stryd values. ● GoldenCheetah CP/ W': We manually added the CP and W' software values to the central table, with all other metrics. ● intervals.icu: Fitness, created a new view covering the full-time range and presenting a Run FTP chart. We couldn't find a way to copy all the values, so just like with the Stryd Powercenter - the weeks in question were recorded. It appears that the platform does not provide a W' value at the time of this page. This gave us 47 CP points, since we don't have a reference model, we averaged all 5 modeled values and compared them to it. The results are in Graphic 1 .

33 Graphic 1. : Comparison of 5 modeled critical power (or FTP) values obtained from different models against the group mean Table 2. visualizes the percentage difference between the five models and the mean value. All values, except the minimum in GC CP, fall within 5% of the mean CP. The discrepancy between the CP/FTP value of the 5 models, for the 47 points we looked at, is between 4W and 32W with a STDEV (Google Sheets function) of 6.59W. Table 2. : Percent difference between the average value between the 5 models and each of the model values (calculated using Google Sheets) CP2 vs AVG Stryd CP vs AV G m F T P v s AV G GC CP vs AVG e F T P v s AV G MIN -2.97% -3.41% -4.75% -6.29% -4.20% MAX 3.06% 2.91% 3.72% 4.38% 4.57% STDEV 1.39% 1.54% 2.09% 2.37% 1.55% For the anaerobic part (W’ and FRC), (Graphic 2.): the values estimated by the models are significantly different - GC W’ is typically about double the WKO5 FRC. Because of this one should target such workouts using the actual power values for a specific duration, rather than modeled based on CP and W’ values. Stryd does not provide W’ value and intervals.icu appears to provide numbers, like those from GoldenCheetah. Graphic 2. : Range of values, provided for the Reserve (“Anaerobic”) abilities of the athlete 3 MINUTES ALL OUT (3MT) The 3-minute full test (3MT) is another way to evaluate 2 CP parameters: critical power and W'. While the original study (Vanhatalo et al., 2007) used cycling, it has also been validated for running (Pettitt et al., 2012). The recommendation of the authors is to use runs that range between 2.5 and 18 minutes. Because the test subject competes with too many 5k, some of these races with a lot of fatigue, some of these races finish earlier, between 7 and 25 minutes and this makes it suitable for such 3MT analysis. Extracting 3MT data can give us an idea of the level of performance decline in non-elite ultrarunners, as we

and 18 minutes. Because the test subject competes with too many 5k, some of these races with a lot of fatigue, some of these races finish earlier, between 7 and 25 minutes and this makes it suitable for such 3MT analysis. Extracting 3MT data can give us an idea of the level of performance decline in non-elite ultrarunners, as we have quite a few overconfident 5k races that have turned into 3MT. From Graphic 3.: Range of values, provided for the Reserve (“Anaerobic”) abilities of the athlete.- (an example 3- minute race type) it is evident - when the race is aggressively paced, chances are your energy reserve (W') is almost depleted and no amount of effort can force it past CP, until the final sprint where the last remaining drops of energy are available.

Description

The study evaluates various Critical Power models using four years of data from a recreational runner.