A client sends you a screenshot before their session. Recovery 41. HRV down nine milliseconds. Six hours twelve minutes of sleep, thirty-eight of it deep. Should they skip the squats?
You have to answer, and you have about four minutes to do it. Answer badly in one direction and you have taught a client that a number on their wrist can cancel a session. Answer badly in the other and you have told someone their $500 watch is rubbish, which is both rude and not quite true.
The useful answer is that a watch measures two or three things well enough to coach off, estimates several more badly enough that acting on them makes decisions worse, and packages the whole lot in a way that invites clients to over-read it. This is how to sort one from the other.
Decide what the device is actually for
A wearable is not a measuring instrument in the way a set of scales is. It is a movement sensor, an optical heart-rate sensor, and a pile of algorithms turning those two signals into things neither sensor can see.
That split is the whole thing. Where the device counts something mechanical and repetitive — steps taken, minutes elapsed, beats per minute during steady work — it does a decent job. Where it infers something physiological and hidden — how many kilojoules you burnt, which stage of sleep you were in, how recovered you are — it is guessing from the same two signals, and the guess carries error the client never sees, because the app shows it as a whole number with a coloured ring around it.
So before you act on any figure from a client's watch, ask two questions. Is the device measuring this or estimating it? And am I looking at one day, or at a direction over several weeks? Almost every bad decision made with wearable data fails one of those two tests.
What to trust, metric by metric
| Metric | How much to trust it | What to do with it |
|---|---|---|
| Steps | Good. A direct count, with a modest under-read | Set a daily floor and coach it like any other habit |
| Heart rate during steady activity | Good | Sanity-check that easy work is actually easy |
| Resting heart rate | Good as that client's own trend, not as a number to compare between people | Watch the direction over weeks; a persistent climb is a doctor's question |
| Calories burnt | Poor. An estimate built from movement and heart rate | Nothing. Never feed it into a calorie target |
| Sleep duration | Reasonable, and the most useful sleep number on the device | Coach the behaviour that produces it — bedtime, not the score |
| Sleep stages (deep, REM, light) | Poor. The device has no access to brain activity | Ignore, and tell the client why |
| HRV | Meaningful as one person's trend on one device. Meaningless as an absolute or across devices | Read the multi-week direction, never a single morning |
| Readiness or recovery score | A proprietary blend of the rows above, weightings undisclosed | Treat as one input alongside how the client says they feel |
Two things are worth naming from that table before going further, because they cause most of the trouble.
The first is that the two metrics coaches most want — calorie burn and recovery — are the two the device is worst at. The second is that clients weight the table upside down. Nobody screenshots their step count. They screenshot the recovery score.
Steps are the one number worth programming around
Steps are the exception that makes wearables worth bothering with at all. The device is counting a real, discrete, repeated event, and a 2024 umbrella review of wearable validation studies published in Sports Medicine found step counts running roughly 9% under a criterion measure, with device errors spread from about 9% under to 12% over.
That is not laboratory-grade, and it does not need to be. A client whose watch under-counts by 9% still walks more this month than last month, and the change is what you are coaching. Consistency of error is worth more here than accuracy of error.
Steps are also the only wearable metric that maps directly onto something you control. Daily activity outside training is the largest movable part of most clients' energy expenditure — the mechanics are in what NEAT is and why it moves the needle — and unlike training volume, you can raise it without adding a session to a week that has no room for one.
Use it like this. Take a fortnight of the client's normal life as a baseline rather than picking a round number, then set a floor a bit above it and hold them to the floor rather than the average. Seven thousand every day beats twelve thousand twice and two thousand five times, because the floor is the thing that survives a bad week. Progress it slowly, the same way you would progress load — the logic is identical to how to apply progressive overload with online clients.
Two cautions. Step counts follow the phone and the watch combined, so a client who leaves both behind produces a low day that is not a real low day. And steps miss cycling, swimming and heavy lifting, so a client whose training is any of those looks sedentary on paper while doing plenty.
Calorie burn is the number most likely to wreck a client's nutrition
Every watch reports calories burnt. No watch measures calories burnt. It infers them from movement, heart rate, and whatever the client entered about their age and weight, and the inference is the weakest thing the device does.
The same 2024 Sports Medicine umbrella review put energy-expenditure error across studies at roughly 21% under to 15% over the criterion measure, and its authors were reluctant to call any device sufficiently accurate for the purpose. For a client eating around 2,000 calories a day, an error band that wide is the difference between a deficit and a surplus.
The practical damage is specific. A client sees "burnt 640 calories" after a spin class, eats those calories back, and cannot understand why the scale has not moved in five weeks. Or their tracking app pulls the watch's figure in as an adjustment and quietly moves their target upward on active days, which is the same problem wearing a nicer interface.
So: never let a wearable's calorie figure set or adjust a nutrition target. Set the target from body weight, activity level and goal, then adjust it from what actually happens to the client's weight and measurements over three or four weeks. That feedback loop is slower than a number on a watch and it is the only one that responds to the client's real physiology.
Say this to the client directly and early, because if you do not, the watch will say it first and it sounds more confident than you do.
Sleep duration is worth coaching. Sleep stages are not
Sleep splits neatly into a number you can use and a number you should retire.
Duration and timing are reasonably tracked, and more importantly they are coachable. A client who is consistently in bed at 11:40 and up at 6:10 has a problem you can work on with them. That is behaviour, and behaviour is your job.
Stages are a different matter. Working out whether someone is in deep, light or REM sleep normally requires brain activity, breathing and eye movement, and a wrist device has none of them — it is inferring architecture from movement and pulse. A 2025 validation study in SLEEP Advances compared six consumer wrist wearables against polysomnography and found agreement on four-stage sleep classification ranging from fair to moderate — Cohen's kappa between 0.21 and 0.53 across the devices tested. The same devices were much better at the simpler question of asleep versus awake.
That is the sentence to give clients, in plain form: your watch has a rough idea of whether you were asleep, and a poor idea of what kind of sleep it was. Chasing a deep-sleep percentage is chasing a number the device is guessing at, and the chasing itself tends to make people sleep worse.
Coach bedtime consistency, wake time, caffeine cut-off, and what happens in the hour before bed. Those move duration, and duration is the part the device gets right.
HRV and recovery scores: a direction, never a verdict
Heart rate variability is real physiology, and it is genuinely informative — with two conditions attached that consumer apps routinely drop.
It is only comparable to itself. Devices sample at different times, over different windows, using different calculations, so one client's 45 and another's 90 say nothing about who is more recovered. Even the same person switching brands starts a new baseline. Treat any absolute HRV figure, and any comparison between two people's figures, as meaningless.
And a single morning tells you almost nothing. HRV moves with alcohol, illness, a late meal, a warm room, a bad night, and the position the arm was lying in. What carries information is a multi-week direction: a client whose baseline has drifted down across a month of hard training while their performance and mood have also fallen is telling you something. One low Tuesday is not.
Readiness and recovery scores are these inputs blended by an algorithm the manufacturer does not publish, weighted in a way you cannot inspect, and presented as a percentage — which is a presentation problem as much as an accuracy one. A percentage implies precision the underlying measurement does not have.
None of this makes the scores useless. It makes them one input among several, and not the deciding one. If you want the data in front of you rather than arriving as screenshots, a client can connect Apple Health or Health Connect once and their steps, sleep, weight and workouts land on their profile — steps and sleep filling in matching habits without anyone logging them by hand.
The client who trains worse because of their recovery score
This is the most common wearable problem in coaching, and it is a coaching problem rather than a data problem.
The pattern: a client checks their score before getting out of bed, sees red, and arrives already convinced their body is not up to it. They warm up expecting to feel bad, drop the top set, cut the last block, and log a session well below what they were capable of. The score was a guess. The lost training was real.
There is no reason to think an unvalidated score outperforms a client's own sense of how they feel once they are warm. So invert the order. The client trains the warm-up and the first working set, then judges. That single change moves the decision from a number produced before any effort to evidence produced during it — which is exactly what RPE gives you with online clients, and it is available to every client whether or not they own a watch.
Give them a rule with a shape, so the decision is already made when they are standing in the gym feeling anxious:
- Score low, first set felt normal. Train the session as written.
- Score low, first set felt genuinely heavy. Drop to the bottom of the prescribed range and finish the session.
- Score low, plus a sore throat, three bad nights, or a real reason. Take the day. The reason is what made the decision, not the score.
- Score high, feeling terrible. Believe yourself, not the ring.
Where a client's scores are consistently poor across weeks and their training is genuinely deteriorating, that is not a session-level decision. That is a block-level one, and it belongs in the planned fatigue management you already have — how periodisation handles accumulating fatigue is the better lever than an ad-hoc rest day every time the app goes red.
For a client who cannot hold that line, turning the score notification off for a month is a legitimate coaching intervention, and their training usually improves.
Working wearable data into a check-in without drowning in it
Pulled into a check-in badly, wearable data doubles the reading and adds nothing. Pulled in well, it answers one question you otherwise have to take the client's word on: what happened between sessions.
Three things are worth looking at, and only three.
Step trend against the last two or three weeks. A client whose daily average has quietly halved explains a stall better than any change in their training will.
Sleep duration, as a pattern. Five nights under six hours is context for a bad week of sessions. One short night is not.
Sessions recorded that you did not program. A client doing two extra spin classes a week is information you want before you decide their programme is not working.
Everything else stays out. If a metric will not change what you write in the reply, reading it is unpaid work — the general principle behind running check-ins that stay useful at scale.
One warning about the reverse direction. Do not rewrite a client's training because of a wearable metric mid-block. A programme is a hypothesis you test over weeks, and swapping sessions in response to daily readings destroys your ability to tell whether it worked — how to write a program that holds together depends on giving the plan long enough to answer the question you set it.
Where your job ends
Wearables now surface things that are not fitness metrics: irregular rhythm notifications, blood oxygen readings, resting heart rate that has climbed and stayed climbed, breathing rate changes, ECG features.
None of these are yours to interpret, and it does not matter how much reading you have done. You are not qualified to tell a client that an irregular rhythm notification is nothing, and you are not qualified to tell them it is something.
The response is the same every time, and it takes one line: that is worth showing your doctor, and in the meantime here is what we are doing with training. Do not speculate about causes, and do not tell them to relax or to panic. Then write down that you referred it.
Two situations deserve particular care. A resting heart rate that has risen and stayed risen over weeks is a medical question, not a deload question — you can deload and refer at the same time, and you should. And a client who is under-eating, over-training and watching their recovery score fall needs help you may not be the right person to give; wearable data does not diagnose disordered eating, but a client's relationship with the numbers can be a signal worth taking seriously.
The policy to give every client
Say this at onboarding, before the watch has had a chance to say anything else:
- Steps we will use. We will set a floor and hold it.
- Sleep duration we will use. Sleep stages we will ignore.
- Calorie burn we will ignore entirely, and we will never eat it back.
- Recovery scores we will use as one input, checked after the warm-up rather than before it.
- Anything that looks like a health alert goes to your doctor, not to me.
Clients accept this readily, because it is a clear position rather than a shrug, and because most already half-suspect the numbers are softer than the interface implies. What they want is someone to tell them which ones to stop worrying about.