Heart rate variability has become the headline metric of the recovery-tracking industry. Wearables report it every morning, apps convert it into a readiness score, and athletes at every level make training decisions based on whether the number is green or red.

It is a genuinely interesting physiological measure. It is also being asked to do a job the evidence doesn't support, and the gap between what it measures and what it's used for is wide enough to cause real problems.

What it actually is

A heartbeat isn't metronomic. The interval between successive beats varies slightly, and that variation reflects the balance between the sympathetic and parasympathetic branches of the autonomic nervous system.

Broadly, higher variability indicates greater parasympathetic influence — a rested, recovered state. Lower variability suggests sympathetic dominance, associated with stress, fatigue, illness or overtraining.

That relationship is real and reasonably well established. The physiology isn't the problem.

Problem one: enormous individual variation

HRV values differ hugely between people. Two equally fit athletes of the same age can have baseline values differing by a factor of two or more, driven by genetics, body composition, and factors nobody fully understands.

This means absolute values are meaningless comparatively. Only within-person change matters, and establishing a reliable personal baseline takes weeks of consistent measurement.

Most people never do that properly, and most consumer devices start producing readiness scores within days.

Problem two: measurement conditions dominate

HRV is extraordinarily sensitive to how and when it's measured.

Time of day matters. Body position matters — lying, sitting and standing produce substantially different values. Breathing rate matters enormously, since respiratory sinus arrhythmia is a major contributor. Whether you've eaten, whether you've had caffeine, ambient temperature, whether you spoke during the measurement.

Get any of these inconsistent and the day-to-day variation from measurement conditions swamps the physiological signal you're trying to detect.

Overnight measurement from a wearable partially addresses this by averaging across a long window in a consistent position, which is why sleep-based HRV is more stable than a spot morning reading. It introduces its own issues — sleep stage composition affects the value, so a night with unusual sleep architecture shifts the number without any change in recovery state.

Problem three: it doesn't distinguish causes

This is the fundamental issue for practical use.

A suppressed HRV reading tells you the autonomic system is in a sympathetically dominant state. It does not tell you why. Candidates include training fatigue, incipient illness, poor sleep, alcohol, psychological stress, dehydration, heat, altitude, or something unrelated to any of these.

The training decision depends entirely on which. Suppressed HRV from accumulated training load might warrant a lighter session. Suppressed HRV from a stressful week might warrant training as normal, since exercise generally helps. Suppressed HRV from developing illness warrants rest.

The number is identical in all three cases. So a readiness score that converts it into a training recommendation is making an inference the data can't support.

Problem four: the signal-to-noise ratio

Day-to-day HRV in a well-trained athlete under stable conditions still fluctuates considerably. Distinguishing a meaningful suppression from ordinary noise requires either a large change or a sustained trend.

Which means single-day readings are close to uninformative, and this is exactly how they're presented and used. The rolling averages that some systems provide are far more useful and far less prominent in the interface, for the obvious reason that a daily number drives daily engagement.

What it's genuinely useful for

Having listed the problems, there are real applications.

Detecting sustained trends. A week or more of consistently suppressed values relative to an established baseline is a meaningful signal worth investigating. Not a diagnosis, a prompt.

Illness detection. HRV suppression often precedes symptomatic illness by a day or two, and this is one of the better-supported applications. Catching a developing infection before an athlete feels ill has genuine value.

Monitoring return from heavy blocks. Watching HRV recover towards baseline after an intensive training period gives a reasonable indication of adaptation.

As one input among several. Combined with subjective wellness, sleep duration, and training load, it contributes to a picture. Alone, it doesn't.

The behavioural harm

One thing I'd flag that gets little attention. Daily readiness scores can produce genuine anxiety in athletes, and there's a documented phenomenon of people sleeping worse because they're worried about their sleep score.

An athlete who sees a red readiness score and concludes they'll perform badly may well perform badly, for reasons entirely unrelated to their autonomic state. That's a self-fulfilling measurement and it's a real cost of putting an ambiguous physiological number in front of somebody every morning with a colour attached.

Several practitioners I've spoken to now deliberately keep the raw data away from athletes, review it themselves, and only raise it when there's a sustained pattern. That seems sensible, and it's the opposite of how the consumer products are designed.