Skip to content
Soccer Science Subscribe

Readiness Scores Can Flag a Change—They Cannot Clear a Player to Train

Ingrid Voss · 5 min read

Wearable readiness scores can flag recovery trends but do not reliably predict soccer performance. Use this four-step check before changing training.

Use a wearable readiness score as a prompt to investigate, not as permission to train hard or an instruction to rest. Sleep, resting heart rate and heart-rate variability (HRV) trends can add useful context. The combined score itself has not been adequately validated as a measure of a soccer player’s ability to perform that day.

For a coach, one unexpectedly low score should usually trigger a short conversation—not an automatic change to the session. Repeated low scores alongside poor sleep, unusual soreness, high fatigue or illness symptoms deserve more attention.

A precise sensor does not make the readiness score accurate

Three separate questions are often collapsed into one:

  1. Did the device measure the signal accurately?
  2. Did its algorithm combine the signals meaningfully?
  3. Does the resulting score predict soccer performance or injury?

The evidence becomes weaker at each step.

Some consumer devices measure resting or overnight heart rate and HRV reasonably well, but accuracy varies by device and setting. In a laboratory comparison involving 53 active adults, WHOOP 3.0 showed close agreement with ECG for overnight HRV, while the other tested devices produced larger individual errors. The study assessed specific, now-older models for one night, so its results should not be transferred automatically to current devices or soccer players (Miller et al., 2022).

Sleep illustrates the same distinction. In that study, Oura Generation 2 and WHOOP 3.0 agreed with polysomnography on sleep-versus-wake classification for 89% and 86% of epochs, respectively. Agreement fell to 61% and 60% when identifying a specific sleep stage. A newer six-device study found that wearables detected more than 90% of sleep epochs but had only 29%–52% specificity for wake, with fair-to-moderate overall sleep-stage agreement (Lee et al., 2025). Total sleep time and longer-term patterns are therefore more defensible coaching inputs than treating reported REM or deep-sleep minutes as laboratory measurements.

The final readiness number adds another uncertain layer. A 2025 evaluation identified 14 composite scores from 10 manufacturers. HRV appeared in 86% of them, resting heart rate in 79%, and physical activity and sleep duration in 71%. Yet no manufacturer disclosed its exact formula, and few supplied peer-reviewed validation of the finished score’s accuracy or relevance (Doherty et al., 2025). Two devices can therefore receive similar raw signals and still return different “readiness” judgments.

What the soccer evidence supports

HRV is physiologically relevant, but it is not a complete readiness test. It reflects autonomic regulation and can respond to training, matches, sleep disruption, psychological stress, illness and other influences. That makes it useful for detecting that something has changed, but often unable to establish what changed or how the player will perform.

A systematic review included 19 studies of adult soccer players. Thirteen reported relationships between linear HRV measures and markers involving performance tests, training load, adaptation, fatigue, recovery or hormones. The studies were of fair average methodological quality and used inconsistent methods, however (Laborde-Cárdenas et al., 2025). This supports HRV as one monitoring input, not as a stand-alone decision rule.

Soccer-specific work also shows why personal baselines matter. During two national-team camps, 34 professional players displayed substantial individual HRV changes after matches. Group analysis detected changes only among the first-choice players and could overlook individual responses (Muñoz-López et al., 2020). Comparing one player with a teammate—or with a generic “good HRV” range—is less useful than comparing that player with their own consistently collected history.

Nor should “recovered” be interpreted as “will play well.” A systematic review and meta-analysis found no significant pooled relationship between HRV and Yo-Yo Intermittent Recovery Test Level 1 performance, although the estimate was imprecise. Most other pooled markers also lacked convincing predictive validity; countermovement-jump height was the notable marker associated with 10 m sprint performance (Duignan et al., 2023).

Even changing training according to HRV is not a guaranteed performance advantage. A meta-analysis found that HRV-guided endurance training improved vagal-related HRV more than predefined training, but its advantages for aerobic fitness and endurance performance were small and statistically non-significant (Düking et al., 2021). Those studies did not test proprietary readiness scores or soccer match performance.

A four-step check before changing training

When a player presents a low score, work through four questions.

1. Is it a real deviation for this player?

Compare the result with the player’s recent trend, not another player or an internet benchmark. Guidance on mobile HRV monitoring favors frequent, consistently collected readings and weekly trends over isolated values. HRV measurements should also be taken while stationary and at rest: movement, poor sensor contact and immediate post-exercise conditions can distort beat-to-beat data (Rogers and Gronwald, 2025).

Check for a loose fit, poor skin contact, a missed night, a changed recording routine or a device switch. Do not compare raw HRV values or scores across brands as if they share a scale.

2. Does the player report the same problem?

Ask briefly about:

  • sleep quality and duration;
  • general fatigue;
  • leg soreness or pain;
  • stress and mood;
  • illness symptoms;
  • confidence about training.

Do not dismiss these answers as less scientific. A systematic review of 56 studies found subjective measures were generally more sensitive and consistent than common objective measures when tracking responses to training load (Saw et al., 2016). In a small study of nine professional soccer players, a four-item wellness index had a better signal-to-noise ratio for post-match fatigue than resting HRV (Rabbani et al., 2019).

3. Does the context explain it?

Review yesterday’s minutes, high-speed work, collisions, travel, heat, hydration, late kickoff and current fixture congestion. Readiness data describe part of the player’s response; they do not replace the record of what the player did.

A low score after a demanding match may be expected. A falling trend during a light week—or a score that conflicts with how the player feels—is a reason to inspect the data and speak with the player.

4. What is the smallest sensible adjustment?

If the score is low but the player feels normal and there are no warning signs, retain the plan while monitoring the warm-up and early work. If a low trend agrees with poor sleep, fatigue and soreness, adjust the dose that creates the most concern: fewer high-speed repetitions, less conditioning volume, longer recovery or a technical role with lower physical load. The aim is not to convert every amber score into a rest day.

Pain, suspected injury, fever, chest symptoms, dizziness or unusual shortness of breath should override the app and prompt appropriate medical assessment. Consumer scores are not return-to-play tests. Oura’s own U.S. Soccer partnership announcement states that its ring is not a medical device and is not intended to diagnose, treat, monitor or prevent illness (U.S. Soccer, 2026).

The useful verdict

Wearable readiness scores are potentially useful for spotting within-player trends, insufficiently validated as finished scores, and not accurate enough to predict a soccer player’s performance from one morning’s number.

If seeing a red or green score appears to influence the player’s answer or confidence, collect the short wellness report before revealing it. Then combine the wearable trend with self-report, recent load and the player’s movement in the warm-up. The device’s best contribution is starting a better question: What changed, and does anything else confirm it?