Abstract
iomechanical assessments of running typically take place inside motion capture labora- tories. However, it is unclear whether data from these in-lab gait assessments are representative of gait during real-world running. This study sought to test how well real-world gait patterns are represented by in-lab gait data in two cohorts of runners equipped with consumer-grade wearable sensors measuring speed, step length, vertical oscillation, stance time, and leg stiffness. Cohort 1 (N= 49) completed an in-lab treadmill run plus five real-world runs of self-selected distances on self-selected courses. Cohort 2 (N= 19) completed a 2.4 km outdoor run on a known course plus five real-world runs of self-selected distances on self-selected courses. The degree to which in-lab gait reflected real-world gait was quantified using univariate overlap and multivariate depth overlap statistics, both for all real-world running and for real-world running on flat, straight segments only. When comparing in-lab and real-world data from the same
2.4 km outdoor run on a known course plus five real-world runs of self-selected distances on self-selected courses. The degree to which in-lab gait reflected real-world gait was quantified using univariate overlap and multivariate depth overlap statistics, both for all real-world running and for real-world running on flat, straight segments only. When comparing in-lab and real-world data from the same subject, univariate overlap ranged from 65.7% (leg stiffness) to 95.2% (speed). When considering all gait metrics together, only 32.5% of real-world data were well-represented by in-lab data from the same subject. Pooling in-lab gait data across multiple subjects led to greater distributional overlap between in-lab and real-world data (depth overlap 89.3–90.3%) due to the broader variability in gait seen across (as opposed to within) subjects. Stratifying real-world running to only include flat, straight segments did not mean- ingfully increase the overlap between in-lab and real-world running (changes of <1%). Individual gait patterns during real-world running, as characterized by consumer-grade wearable sensors, are not well-represented by the same runner’s in-lab data. Researchers and clinicians should consider “borrowing” information from a pool of many runners to predict individual gait behavior when using biomechanical data to make clinical or sports performance decisions. Keywords:wearable technology; depth statistics; unsupervised learning; free-living gait; biomechanics 1. Introduction Individual differences in running biomechanics have been associated with both injury and performance [1–4]. Running gait has traditionally been assessed during in-lab motion capture sessions, but data collected under these conditions are only useful for real-world clinical and sporting applications if the in-lab data are well-representative of gait patterns adopted during real-world training. Since virtually all training takes place outside of the lab, the generalizability of models or inferences based on in-lab biomechanical data to real- world running is of paramount importance for both clinicians and researchers. Likewise, generalizability of in-lab data presents challenges in professional sport, for similar reasons: Sensors2024,24, 2892.
Sensors2024,24, 2892 2 of 22 gait characteristics in matches or competitions may not reflect those seen during in-lab evaluations conducted as part of training or injury rehabilitation. While a recent systematic review found only minor biomechanical changes when comparing gait patterns during treadmill versus overground running [5], other work that more directly compares in-lab to real-world running has noted changes in various aspects of running gait. Lafferty et al. report that video-based gait analysis showed differences in gait variables including footstrike angle, tibial inclination, and pelvic drop when com- paring indoor treadmill versus outdoor track running [6], and Benson et al. developed a classifier based on sensor-measured gait features that could differentiate between tread- mill and sidewalk running with ~80% accuracy, suggesting distinctive differences in the characteristics of gait in the lab versus in the real world [7]. To date, though, the previous literature has focused on differences in the mean value of individual gait metrics, and has not considered how overall gait patterns are distributed during in-lab and real-world running. Moreover, whether differences seen in real-world running can be ascribed to changes in the environment (e.g., turns, inclines, declines) remains unclear. Consumer-grade sensors are particularly attractive for real-world gait assessment because of their low cost, wide usage, and ability to synchronize data with cloud-based training platforms, which allows researchers and clinicians to remotely collect and monitor gait data on hundreds or thousands of runners at once [8]. Recent work has explored using both research-grade and consumer-grade wearable sensors to characterize gait patterns during real-world running, due to the ease with which wearable sensors can be used outside of the lab [9,10]. Measuring a runner’s full gait pattern during real-world running is challenging despite the utility of wearable sensors. Both consumer-grade and research-grade wearable sensors measure only a limited number of gait metrics compared with what is possible with in-lab motion capture equipment, and not all devices measure the same gait metrics. However, using multiple devices together can capture gait metrics such as speed, stride length, vertical oscillation, ground contact time, and leg stiffness. Many of these same gait metrics
consumer-grade and research-grade wearable sensors measure only a limited number of gait metrics compared with what is possible with in-lab motion capture equipment, and not all devices measure the same gait metrics. However, using multiple devices together can capture gait metrics such as speed, stride length, vertical oscillation, ground contact time, and leg stiffness. Many of these same gait metrics are used in simplified biomechanical models of running, such as the well-studied mass- spring model, which explains many key aspects of running biomechanics [11]. Combining these gait metrics gives rise to the idea of a “gait pattern”—a set of gait metrics that jointly represent the body’s movement. Comparing sensor-measured gait patterns between runners or between different conditions (e.g., in-lab versus real-world) is a straightforward way to quantify similarities or differences in gait. To this end, the primary goal of this study was to compare the distribution of gait patterns during in-lab and real-world running. Three related questions are relevant when considering whether gait patterns during in-lab running are representative of gait patterns during real-world running. First, is a runner’s gait pattern during an in-lab gait analysis a good representation of that same runner’s real-world gait pattern? This question is relevant to in-lab biome- chanical analyses and in-lab gait retraining interventions, which are done with the aim of generalizing from in-lab running to real-world running in the same individual. Second, is a set of in-lab gait data from a large pool of runners a good representation of the real-world gait pattern that might be observed in a new runner from the same population? This question is relevant when constructing predictive models based on in-lab data that aim to generalize to real-world running data from new, unseen runners (i.e., runners whose data were not used to develop the model). For example, Matijevich et al. [12] developed a sensor-based model for predicting compressive forces on the tibia. Successfully applying this model to free-living data from runners in the same population would require the in-lab data collected on the subjects who formed the “training set” to be a good match for the real-world running
were not used to develop the model). For example, Matijevich et al. [12] developed a sensor-based model for predicting compressive forces on the tibia. Successfully applying this model to free-living data from runners in the same population would require the in-lab data collected on the subjects who formed the “training set” to be a good match for the real-world running data from a new, unseen “test” subject. Third, is a set of in-lab gait data from a large pool of runners a good representation of the real-world gait pattern that might be observed from a new runner from a new population, potentially in a different geographic location? This question is relevant when
Sensors2024,24, 2892 3 of 22 discussing the translatability of study findings, i.e., whether statistical inferences or predic- tive model performance from a study on one population of runners (e.g., healthy adults in one location) will generalize to a different or more specialized population of runners (e.g., college-aged females in a different location). For example, this question would be important for clinicians who want to apply findings from a published study to the real-world training of a patient from a new population, and for researchers who want to apply a published predictive model to a new sample of runners. This study addressed each of these questions by quantifying the degree of overlap between gait patterns during in-lab and real-world running, as measured by a set of consumer-grade wearable sensors. Further, this study disaggregated the effects of the real-world running environment (inclines, declines, and turns) from changes in gait pattern on flat, straight settings by comparing gait patterns during all real-world running, versus real-world running only on flat, straight segments. 2. Materials and Methods 2.1. Overview This study involved two separate cohorts of runners representing different populations of potential interest to researchers and clinicians. Cohort 1 consisted of healthy male and female runners aged 18 and older who completed an in-lab treadmill run while equipped with a set of consumer-grade wearable sensors. These same runners completed five real- world, free-living runs using the same set of sensors. Cohort 2 followed a different protocol, which was designed to assess the generalizability of the findings from Cohort 1 to a new population, as well as to determine the potential sources of gait differences between in-lab and real-world running, specifically, the influence of turns, inclines, and declines. Cohort 2 consisted of healthy female runners aged 18 and older in a different geographic location who completed a 2.4 km run on a measured course with known segments of flat, turning, incline, and decline running while wearing the same set of sensors as Cohort 1. This cohort also completed five real-world, free-living runs, again using the same set of sensors. Gait metrics from the wearable sensors were
18 and older in a different geographic location who completed a 2.4 km run on a measured course with known segments of flat, turning, incline, and decline running while wearing the same set of sensors as Cohort 1. This cohort also completed five real-world, free-living runs, again using the same set of sensors. Gait metrics from the wearable sensors were used to characterize the gait pattern for each runner, and the distributions of these metrics during in-lab and real-world run- ning were compared to quantify the proportion of overlap in gait patterns across these distributions. 2.2. Participants Cohort 1.The inclusion criteria for Cohort 1 were designed to capture a pool of runners representative of the broader population of runners. Healthy runners aged 18 and older were recruited, with no upper limit on age. Participants were required to run at least three times per week with at least one run of 40 min or longer, were required to have no current musculoskeletal injury that prevented them from doing their usual running training, and were required to meet American College of Sports Medicine preparticipation guidelines for exercise [13]. Runners were recruited from the community via social media, flyers at local running stores, and in person recruitment at local running events. Recruitment and data collection for Cohort 1 took place in Greenville, North Carolina, which is located in a region with predominantly flat terrain. All participants provided written informed consent, and the study was approved by the East Carolina University and Medical Center IRB and the Indiana University IRB (protocols # 21-001137 and 12040). The sample size for Cohort 1 was determined via a learning curve power analysis for a predictive modeling goal detailed elsewhere [14], which indicated that a minimum of 40 participants were needed. Cohort 2.The inclusion criteria for Cohort 2 were designed to construct a more homogenous and specialized population of runners to assess the generalizability of in-lab data to a new population of athletes. One such specialized population often studied in prospective research on running injuries is young adult female runners, who may be at greater risk of overuse
participants were needed. Cohort 2.The inclusion criteria for Cohort 2 were designed to construct a more homogenous and specialized population of runners to assess the generalizability of in-lab data to a new population of athletes. One such specialized population often studied in prospective research on running injuries is young adult female runners, who may be at greater risk of overuse injury (e.g., Davis et al., Rauh et al. [15,16]). In service of this goal of testing the generalizability of findings from the in-lab data, Cohort 2 included women
Sensors2024,24, 2892 4 of 22 aged 18–32 were recruited who fulfilled the same inclusion criteria as Cohort 1 (running at least three times per week with one run lasting at least 40 min, and no current injuries or contraindications for exercise). Runners were recruited from students at a large university via flyers on campus and at a local running club. Recruitment and data collection for Cohort 2 took place in Bloomington, Indiana, which is located in a region with predominantly hilly terrain. All participants provided written informed consent, and the study was approved by the Indiana University IRB (protocol #17923). The sample size for Cohort 2 was designed to recruit a similar number of female subjects as were recruited for Cohort 1. 2.3. Wearable Sensors and Gait Metrics Three consumer-grade wearable sensors were used to collect gait data during in-lab and real-world running: a sports watch with global navigation satellite system (GNSS) capabilities (Garmin Forerunner 245, Garmin Ltd., Olathe, KS, USA) worn on the partici- pant’s left wrist; a chest strap heart rate monitor with an integrated accelerometer (Garmin HRM-Run and HRM-Tri, Garmin Ltd., Olathe, KS, USA), which was worn around the chest, centered over the heart and inferior to the sternum, and a foot pod with an inte- grated inertial measurement unit (Stryd v2, Stryd Inc., Boulder, CO, USA), which was placed on the distal shoelaces of the left shoe. This combination of devices was chosen because these devices are already in wide use, record and synchronize their data to remote cloud-based training platforms, and capture key biomechanical aspects of gait that can be used to characterize a runner’s gait pattern. Six separate matched sets of these three devices were used to reduce any device-specific systematic errors, and to enable parallel enrollment of multiple subjects. Three of these device sets used the HRM-Run model of chest strap sensor, and three of these devices sets used the HRM-Tri model of chest strap sensor; using multiple variants of this device (both of which are in wide use) expanded the real-world generalizability of predictive models built as a separate part of the
enable parallel enrollment of multiple subjects. Three of these device sets used the HRM-Run model of chest strap sensor, and three of these devices sets used the HRM-Tri model of chest strap sensor; using multiple variants of this device (both of which are in wide use) expanded the real-world generalizability of predictive models built as a separate part of the project [14]. Validation testing on a separate cohort of ten runners showed that the chest strap-measured gait metrics show close agreement in the gait metrics measured across the two variants of the chest strap (see Supplementary Data S1, which details mean absolute percentage differences between devices). Five sensor-measured gait metrics were selected to represent a runner’s gait pattern: running speed, step length, vertical oscillation, stance time, and leg stiffness. These five specific gait metrics were selected because they correspond to key parameters of the mass- spring model of running, a simple and well-studied model that describes numerous aspects of running gait [11], and because these specific metrics are measured with acceptable accuracy by the wearable sensors. Step length and vertical oscillation were measured with the chest strap, while stance time and leg stiffness were measured by the foot pod. Since the GNSS technology of the sports watch does not work indoors, speed data from the foot pod was used to measure running speed on the treadmill. Both the foot pod and GNSS speed estimates have errors of <2% compared to ground-truth running speed in previous research, and the foot pod’s speed data showed no statistically significant bias against the ground-truth treadmill speed (see Supplementary Table S1) [17,18]. The accuracy of the individual gait metrics was determined empirically for runners in Cohort 1 using motion capture data collected during the in-lab treadmill run (See Supple- mentary Table S1 for full device metric validation results including correlation coefficients and Bland–Altman limits of agreement). Though not all metrics are measured by the devices with equal absolute accuracy, this study’s validation and other validation studies on the same devices have determined that these metrics are measured with sufficiently high accuracy compared to in-lab metrics
treadmill run (See Supple- mentary Table S1 for full device metric validation results including correlation coefficients and Bland–Altman limits of agreement). Though not all metrics are measured by the devices with equal absolute accuracy, this study’s validation and other validation studies on the same devices have determined that these metrics are measured with sufficiently high accuracy compared to in-lab metrics to detect changes in gait within and across individuals [19–21]. In the case of gait metrics measured by two sensors, the sensor which measured that gait metric more accurately was used—this was the chest strap for vertical oscillation, and the foot pod for stance time. Since speed, cadence, and stride length are mathematically linked, only speed and stride length were used in the representation of a runner’s gait
Sensors2024,24, 2892 5 of 22 pattern, because (1) runners often vary their speed by changes in stride length as opposed to cadence [22,23], and (2) the devices quantize cadence by mapping it to an integer number of strides (foot pod) or an integer number of steps (chest strap) per minute. This quantization process introduces errors compared to using stride length, which is measured by the chest strap to the millimeter. During all runs, the chest strap and foot pod streamed their data wirelessly to the sports watch, which recorded the gait metric values once per second alongside GNSS- determined latitude, longitude, and speed (during outdoor running). All gait metrics for each running session were saved in a single Flexible and Interoperable Data Transfer (FIT) protocol file. 2.4. Protocol Cohort 1, in-lab run.Participants in Cohort 1 first completed a 38 min in-lab treadmill run at speeds ranging from 30% slower to 25% faster than each runner’s self-reported preferred running speed for a “typical training run.” This range of speeds was designed to increase the variability in each runner’s gait as observed in the lab, as most gait-related parameters change as a function of speed. The range of speeds was selected by comparing data on self-reported preferred running speed from a previous in-lab study [24] with known values for the typical walk–run transition speed in healthy adults [25] and predictive equations for estimating lactate threshold from training pace [26]. The range of 30% slower to 25% faster kept the slowest speeds above the walk–run transition for most adults, avoiding uncomfortably slow speeds, and kept the fastest speeds below each runner’s predicted lactate threshold, avoiding excessive fatigue. The speeds were presented in a semirandomized fashion, with the slowest two speeds first, then a block of randomized speeds, followed by the fastest two speeds at the end. This ordering was chosen to maximize the range of speeds covered by each subject and to minimize early-onset fatigue that would prevent subjects from completing the protocol. To minimize any potential order effects, each subject was randomly assigned one of four randomized block orders (speed ordering for
Description
This study tests the representativeness of in-lab gait data for real-world running.