A wearable presents dozens of figures, and only a handful of them are measured. Distinguishing the primitive signals from the estimates explains most of what goes wrong with the output.

The primitive signals are few

A typical unit contains an accelerometer, a gyroscope, a magnetometer, a positioning receiver and often an optical or electrical sensor for heart activity. Everything reported is built from those inputs.

Each measures something narrow and physical, such as acceleration along an axis or the interval between heartbeats. None of them measures fatigue, effort or readiness.

Every concept a coach actually cares about is therefore inferred rather than observed.

Inference introduces assumptions at each step

Turning acceleration into distance requires integrating over time, and small errors accumulate rapidly through that process. Positioning data is used to constrain the drift, which works well outdoors and poorly under a roof.

Turning heart intervals into an estimate of internal load requires assumptions about an individual's maximum and resting values, both of which change with fitness and with the day.

The chain from signal to insight is short in presentation and long in assumption, and each assumption is a place where the estimate can fail.

Sleep and recovery figures are the furthest from measurement

Devices infer sleep stages from movement and heart activity, which are indirect proxies for a process defined by brain activity. The inference is reasonable on average and unreliable for a given night.

Recovery scores combine those estimates with training data through another proprietary formula, adding a second layer of inference on top of the first.

These outputs are useful for spotting trends across weeks and are not diagnostic of anything, and any concern about a player's health is a matter for qualified clinical staff.

Validation is uneven across the outputs

Positional accuracy and heart interval measurement have been compared extensively against laboratory reference methods, so their error ranges are reasonably well characterised.

Composite scores are validated far less often, partly because there is no reference standard to validate them against.

A product can therefore be accurate in its measurements and unproven in its conclusions at the same time, which is the normal situation rather than an unusual one.

Knowing the layer changes how the number is used

Staff who understand which figures are measured treat those as evidence and treat composite scores as prompts for a conversation with the player.

That distinction is why experienced practitioners ask about sleep and soreness rather than reading a readiness figure aloud.

The wearable is best understood as an instrument that narrows where to look, and the looking is still done by people.