← Back to library
article 2025 17 pages

Limited Interchangeability of Smartwatches and Lace-Mounted IMUs for Running Gait Analysis

Theodor Meingast, Bryson Carrier, Amanda Melvin, Kenneth M. Kozloff, Alexandra F. DeJong Lempke, Adam S. Lepley

Journal
Sensors
DOI
10.3390/s25175553
Population
physically active adults
View on DOI ↗

Abstract

patiotemporal running metrics such as cadence, stride length (SL), and ground contact time (GCT) are important for assessing performance and injury risk. However, such metrics are traditionally assessed using laboratory-based tools that are often inaccessible in applied settings. Wearable devices including smartwatches and lace-mounted inertial measure- ment units (IMUs) offer promising alternatives, yet cross-device agreement in reporting spatiotemporal variables remains unclear. This study evaluated agreement between a commercial smartwatch and lace-mounted IMUs across varied distances and environ- ments in 65 physically active adults (33 female/32 male, height: 171.0±8.9 cm; weight: 70.9±15.2 kg). Participants completed indoor and outdoor runs (2.5 km, 5 km, 10 km, 20 km) wearing both devices simultaneously. Average cadence demonstrated acceptable agreement (MAPE = 4.1%, CCC = 0.66) and supported equivalence, particularly among males, during outdoor conditions, and longer run distances. In contrast, peak cadence showed

65 physically active adults (33 female/32 male, height: 171.0±8.9 cm; weight: 70.9±15.2 kg). Participants completed indoor and outdoor runs (2.5 km, 5 km, 10 km, 20 km) wearing both devices simultaneously. Average cadence demonstrated acceptable agreement (MAPE = 4.1%, CCC = 0.66) and supported equivalence, particularly among males, during outdoor conditions, and longer run distances. In contrast, peak cadence showed weak correlation (MAPE = 5.3%, CCC = 0.29), and SL and GCT demonstrated poor agreement (MAPE = 14.9–19.0%, CCC = 0.30–0.39) across all conditions. While average cadence may serve as a metric for cross-device comparisons, especially for males, and longer-distance outdoor runs, other spatiotemporal metrics demonstrated poor agreement, limiting interchangeability. Understanding device-specific capabilities is essential when interpreting wearable-derived gait data. Further validation using gold-standard tools is needed to support accurate and applied use of wearable technologies. Keywords:wearable sensors; biomechanics; gait; field-based assessment; biometric technology; fitness tracker; activity monitor 1. Introduction Spatiotemporal running biomechanics, including ground contact time (GCT), cadence, and stride length (SL), play a crucial role in assessing running performance and injury risk. These metrics have been shown to serve as key indicators of fatigue, and are often used as indicators of biomechanical efficiency, particularly during prolonged endurance events [1–3]. For example, shorter GCT, increased cadence, and optimized SL have been associated with improved endurance performance [4,5]. Additionally, alterations in spatiotemporal variables have been linked to injury susceptibility. Increased GCT, longer SL, and decreased cadence have been associated with the development of exercise-related lower leg pain, while decreased Sensors2025,25, 5553 https://doi.org/10.3390/s25175553

Sensors2025,25, 5553 2 of 17 cadence and increased SL have been identified as risk factors for bone stress injuries [6–10]. Given the key role of spatiotemporal parameters in running biomechanical profiles, ensuring accessible and accurate measurement of these variables is crucial for individuals seeking to assess running performance and running-related injury. The criterion standard for measuring spatiotemporal variables requires 3-dimensional (3-D) motion capture systems, force plates, and high-speed video analysis in controlled laboratory environments. While these methods provide good accuracy and reliability, these systems are cost-prohibitive, time-intensive, and largely inaccessible to recreational runners and clinicians who wish to monitor performance and injury risk in outdoor set- tings. This limitation has driven interest in wearable technology as an alternative for field-based running assessments [11]. Advancements in wearable sensor technologies, such as improvements in sampling rates, integration of inertial measurement units (IMU) with global positioning systems (GPS) and other sensors, and unique device placements, have enabled runners to monitor biomechanical data in real time. Among commercially available devices, lace-mounted IMUs can provide detailed kinetic, kinematic, and spa- tiotemporal variables during running. Previous studies [12,13] have validated the accuracy of the lace-mounted RunScribe TM IMU sensors against laboratory-based standards, re- porting minimal mean differences and excellent agreement for spatiotemporal variables (cadence: ICC = 0.97, r = 0.93; SL: ICC = 0.80–0.86, MAE = 0.7–0.8 m; GCT: ICC = 0.92–0.93, MAE = 27–29 ms) [12,14,15]. Additionally, these devices have demonstrated sensitivity in determining changes in spatiotemporal biomechanics in response to varying speeds, surfaces, and intentional spatiotemporal gait-training modifications [2,16]. These findings suggest that lace-mounted IMUs serve as an accurate and practical tool for consumer use in monitoring running biomechanics. Conversely, smartwatches offer a more convenient, affordable, and widely adopted wearable device that could serve as an alternative for tracking running biomechanics. How- ever, smartwatch validation in accurately measuring spatiotemporal variables is limited and has yielded mixed findings. Some studies suggest smartwatches can detect variability in gait patterns but have demonstrated mixed error and agreement statistics when estimat- ing SL during gait using Garmin, Apple and Samsung smartwatch devices (RSME = 5.29 cm; MAE =

as an alternative for tracking running biomechanics. How- ever, smartwatch validation in accurately measuring spatiotemporal variables is limited and has yielded mixed findings. Some studies suggest smartwatches can detect variability in gait patterns but have demonstrated mixed error and agreement statistics when estimat- ing SL during gait using Garmin, Apple and Samsung smartwatch devices (RSME = 5.29 cm; MAE = 0.13 m; ICC = 0.60) [17–19]. Additionally, the ability to monitor dynamic changes in running biomechanics remains uncertain. For example, an investigation using a Garmin smartwatch was not able to detect alterations in spatiotemporal variables when partici- pants intentionally modified their running styles, raising questions about measurement sensitivity [20]. Despite the increasing availability of wearable devices for runners, no study, to our knowledge, has directly assessed the agreement between lace-mounted IMUs and smart- watches in reporting spatiotemporal variables during running activities. The level of agreement between consumer wearable devices has direct implications for applied use. Establishing whether two commonly used consumer technologies, positioned at different anatomical sites, provide comparable data is essential, as discrepancies could influence how athletes, coaches, and clinicians interpret performance or injury risk. A clear understanding of device agreement helps ensure that training and rehabilitation decisions are based on accurate, reliable information. Therefore, the purpose of this study was to evaluate the agreement between lace-mounted IMUs and a commercially available smartwatch for measuring spatiotemporal variables during running across different distances and envi- ronments. Establishing the level of agreement between these devices could improve our understanding of the interchangeability, accuracy, and real-world application of wearable devices in running assessments.

Sensors2025,25, 5553 3 of 17 2. Materials and Methods The data presented in this study is a subset of an overall dataset focused on examining wearable technology in physically active populations. Adult participants over the age of 18 were recruited from a university population and surrounding community and were included if they self-reported participating in moderate to vigorous physical activity at least three days per week. Participants were excluded if they had a contraindication to intense exercise (e.g., cardiovascular disease, significant musculoskeletal or neurological impairments, etc.) or were pregnant at the time of testing. All participants provided written informed consent prior to testing, and all procedures were approved by the University’s Institutional Review Board (IRB#: HUM00220366). All participant demographic information was collected at baseline through electronic surveys (REDCap, Vanderbilt University, Nashville, TN, USA). Participants then completed three testing sessions, each separated by one to two weeks (days between visit one and two: 11.0±5.4; between visit two and three: 11.4±6.7). Each participant completed an outdoor 5 km run at visit one, an indoor 5 km run at visit two, and was assigned to either an indoor or outdoor run of variable distance (2.5 km, 10 km, 20 km) for their third visit. Cohort determinations for visit three were based on self-reported running experience (days per week engaged in running activity) and self-reported ability to complete the distance at the time of initial study participation screening. All participants were instructed to complete each run at a self-selected light to somewhat hard pace (RPE 11–13) using the Borg Rating of Perceived Exertion scale [21]. During all trials, participants wore a smartwatch (Apple Watch series 7, watchOS 9, Apple Inc., Cupertino, CA, USA, sampling frequency: not reported) on their left wrist, and bilateral lace-mounted IMUs (RunScribe Pods, Scribe Labs, Moss Beach, CA, USA, sampling frequency: 250 Hz) secured using the sensor-specific lace cradles at the midfoot on their running shoes (Figure). This study was designed as a direct comparison between the two devices to evaluate their agreement in measuring spatiotemporal running variables. Importantly, we did not aim to validate these devices against a criterion

IMUs (RunScribe Pods, Scribe Labs, Moss Beach, CA, USA, sampling frequency: 250 Hz) secured using the sensor-specific lace cradles at the midfoot on their running shoes (Figure). This study was designed as a direct comparison between the two devices to evaluate their agreement in measuring spatiotemporal running variables. Importantly, we did not aim to validate these devices against a criterion measure (e.g., 3D motion capture or force plates); therefore, the results should be interpreted solely as a comparative analysis of cross-device agreement. Investigators ensured proper fit of each device per manufacturer instructions so that the devices did not move excessively during activity. Investigators reset device settings and created individual profiles using the profile settings for each device prior to each assessment, including participant date of birth, sex, height, and weight. The lace-mounted IMUs were additionally calibrated to positioning on the foot immediately prior to each running trial. All indoor trials were completed on a motorized treadmill (4Front, Woodway, Waukesha, WI, USA). All outdoor runs were completed on pre-determined GPS-measured routes with standardized start and stop locations and consisted of loops around a university campus that aimed for continuous running and limited road intersections. All outdoor runs were completed on concrete sidewalks, with elevation gains of approximately 7 m for the 2.5 km and 5 km courses and approximately 105 m for the 10 km and 20 km courses. Prior to the outdoor runs, participants were made familiar with the route and were remotely monitored by lab staff during their run via a tracking device (AirTag, Apple Inc., Cupertino, CA, USA). Spatiotemporal variables across all runs, including average GCT (ms), peak and average cadence (steps/minute), and average SL (m), were exported from each device’s exercise recording app following each trial. Per developer documentation [22–24], SL is reported differently between devices, as the smartwatch reports stride length as step length, or half of the defined SL. Thus, smartwatch SL values were corrected (eq. smartwatch SL = reported step length×2) to account for this difference.

developer documentation [22–24], SL is reported differently between devices, as the smartwatch reports stride length as step length, or half of the defined SL. Thus, smartwatch SL values were corrected (eq. smartwatch SL = reported step length×2) to account for this difference.

Sensors2025,25, 5553 4 of 17 Figure 1.Image showing placement of bilateral lace-mounted IMUs, secured with the specific lace cradles at the midfoot of the running shoes. Data Analysis For all trials, overall agreement was evaluated via tests of error, linearity, and equiva- lence between the lace-mounted IMUs and smartwatch for each variable. Mean absolute error (MAE) and mean absolute percentage error (MAPE) were calculated for error analysis. Linearity was established via Lin’s Concordance Correlation Coefficient (CCC), Pearson’s Product Moment Correlation (r), and Deming Regression. Correlation coefficients were interpreted as follows: 0 to <0.2, very weak;≥0.2 to <0.4, weak;≥0.4 to <0.6, moderate; ≥0.6 to <0.8, strong; and≥0.8 to 1.0, very strong [25]. Equivalence testing was performed via confidence interval for difference in means, using the 90% confidence interval from paired t-tests and 10% (±5%) of the criterion mean as the equivalence window [26]. A 10% equivalence window was selected to align with prior wearable validation research where it is widely applied for cross-study comparisons, and because 5–10% changes in spatiotemporal variables, specifically step rate, are known to produce clinically meaningful alterations in lower-extremity loading and running kinematics [27–30]. Binary results for equivalence testing are presented as “Supported” or “Not Supported.” In addition, com- bined agreement criteria were set at MAPE < 10%, CCC > 0.7, and equivalence supported at 10% (±5%) of the criterion mean for the equivalence window, based on the 90% CI [31]. Data were further stratified by sex, distance, and environment for additional analyses (Tables–3). Environment was categorized by indoor and outdoor trials and distance presented as short (2.5 km, 5 km) and long-distance trials (10 km, 20 km). Supplemental stratifications were also conducted on height (<166 cm, 166–175 cm, >175 cm) and weight (<60 kg, 60–70 kg, 70–80 kg, 80–90 kg, >90 kg) 3. Results A total of 65 participants (33 female/32 male, height: 171.0±8.9 cm; body mass: 70.9±15.2 kg) were included. There was a total of 192 running trials that were captured by the smartwatch and 191 by the lace-mounted IMU sensors and included in this study. Dis- crepancies in sample size across analyses reflect instances

70–80 kg, 80–90 kg, >90 kg) 3. Results A total of 65 participants (33 female/32 male, height: 171.0±8.9 cm; body mass: 70.9±15.2 kg) were included. There was a total of 192 running trials that were captured by the smartwatch and 191 by the lace-mounted IMU sensors and included in this study. Dis- crepancies in sample size across analyses reflect instances where one device did not report a given variable, and these trials were therefore excluded from that pairwise comparison. The number of runs included for each analysis can be found in Tables–3.

Sensors2025,25, 5553 5 of 17 Overall device agreement for average cadence between the smartwatch and lace- mounted IMU sensors demonstrated less than 10% error between devices (MAPE: 4.1%). Correlation coefficients were moderate (r = 0.74; CCC = 0.66), and equivalency testing was supported (Table). Stratifications for average cadence revealed that agreement between devices was greater for males, longer distances, and outdoor trials, as these analyses met accuracy thresholds of MAPE < 10%, CCC > 0.7, and equivalence supported at 10% (±5%) of the criterion mean (Tables–3). There was mixed agreement between devices for peak cadence. Although measure- ment error remained low for overall data (MAPE < 10% across all stratifications), correlation values were considered weak (r = 0.45; CCC = 0.29), and equivalence was not supported (Table). Sex, distance, and environment stratifications did not influence accuracy statistics (Tables–3). There was poor agreement overall for SL (corrected SL for smartwatch) and GCT metrics, with poor error rates (MAPE range = 14.9–19.0%), weak-to-moderate correlation coefficients (r range = 0.51–0.61; CCC = 0.30–0.39), and unsupported equivalence testing (Table). Sex, distance, and environment did not influence interpretation of agreement statistics for SL and GCT (Tables–3). Equivalence plots and supplemental stratifications for height and weight can be found in Appendix–A6. Table 1.Agreement statistics between lace-mounted IMUs and smartwatch for spatiotemporal variables for all data and gender stratified comparisons. n Device Average (SD) MAE MAPE r CCC Deming Inter- cept/Slope Overall * AC (step /min) 191 Lace-mounted IMU 169.1 (13.3) — — — — — 192 Smartwatch 162.8 (16.2) 6.8 4.1% 0.74 0.66 −72.4/1.39PC (step/min) 191 Lace-mounted IMU 198.7 (25.0) — — — — — 190 Smartwatch 178.7 (15.9) 20.5 9.5% 0.45 0.29 95.2/0.42 SL (m) 190 Lace-mounted IMU 2.12 (0.39) — — — — — 190 Smartwatch 1.90 (0.28) 0.32 14.4% 0.51 0.39 0.86/0.49 GCT (ms) 191 Lace-mounted IMU 318.1 (69.9) — — — — — 190 Smartwatch 254.5 (32.1) 66.6 19.0% 0.67 0.30 143.59/0.35

190 Smartwatch 1.90 (0.28) 0.32 14.4% 0.51 0.39 0.86/0.49 GCT (ms) 191 Lace-mounted IMU 318.1 (69.9) — — — — — 190 Smartwatch 254.5 (32.1) 66.6 19.0% 0.67 0.30 143.59/0.35

Sensors2025,25, 5553 6 of 17 Table 1.Cont. n Device Average (SD) MAE MAPE r CCC Deming Inter- cept/Slope FemaleAC (step/min) 86 Lace-mounted IMU 168.3 (13.9) — — — — — 88 Smartwatch 161.4 (17.6) 7.7 4.6% 0.65 0.58 −92.7/1.51 PC (step/min) 86 Lace-mounted IMU 197.3 (23.5) — — — — — 87 Smartwatch 178.4 (12.2) 18.3 8.6% 0.43 0.25 116.8/0.32 SL (m) 86 Lace-mounted IMU 1.99 (0.32) — — — — — 87 Smartwatch 1.86 (0.28) 0.28 13.4% 0.34 0.31 0.57/0.65 GCT (ms) 86 Lace-mounted IMU 330.2 (74.3) — — — — — 87 Smartwatch 255.8 (30.6) 78.3 21.4% 0.61 0.23 159.3/0.29 * AC (step /min) 102 Lace-mounted IMU 169.5 (12.8) — — — — — Male 101 Smartwatch 163.6 (15.4) 6.1 3.7% 0.82 0.74 -57.0/1.3 PC (step/min) 102 Lace-mounted IMU 198.9 (25.4) — — — — — 100 Smartwatch 178.8 (18.7) 21.4 9.9% 0.47 0.33 70.0/0.55 SL (m) 101 Lace-mounted IMU 2.25 (0.40) — — — — — 100 Smartwatch 1.95 (0.27) 0.36 15.5% 0.57 0.37 0.92/0.46

Description

The study assesses the interchangeability of smartwatches and IMUs for running metrics.