Most people who own a sleep tracker have had some version of the same experience: you wake up feeling fine, look at the app, see a 61 — and feel worse. The number arrived before the self-assessment did, and it won.
The data isn’t useless. But its usefulness is narrower than the apps suggest, and the ways sleep tracking is most often used are the ways it’s least accurate. What follows is a clear-eyed account of what consumer trackers actually measure, where the research says they mislead, and what the one signal worth watching actually is.
What a consumer tracker actually measures
Most wrist-based sleep devices — Fitbit, Oura, Apple Watch, Garmin — use two primary signals: actigraphy, which detects movement, and photoplethysmography, which measures heart rate and heart rate variability through the skin. From those signals, the algorithm infers when you were asleep, when you were in light NREM sleep, when in slow-wave sleep, when in REM.
The clinical method for measuring sleep stages is polysomnography: electrodes on the scalp reading brain electrical activity, electrodes near the eyes tracking eye movements, sensors on the chin measuring muscle tone. That’s how researchers established the architecture of sleep in the first place — the 90-minute cycles, the progression through stages described in sleep stages, the way deep sleep concentrates early and REM accumulates later in the night.
A consumer tracker cannot read brain waves from a wrist. What it does instead is apply statistical correlations — between HRV patterns and known sleep-stage signatures in research populations — to your nightly data. When those correlations hold, the estimate is in the right neighborhood. When they don’t, the number is confident and wrong.
Population-level statistics are also what make individual nights harder to interpret. When a tracker’s algorithm says users average 90 minutes of deep sleep, it’s working with thousands of data points and normal distributions. When it applies the same statistical model to your specific Tuesday night — one data point, from one person, with one particular body and one particular day behind it — the precision of that “67 minutes of deep sleep” figure is far lower than the decimal place implies. Several studies comparing wrist-based devices to laboratory polysomnography have found reasonable agreement on total sleep time (typically within 20–30 minutes) but substantially lower agreement on sleep stage breakdown, especially the distinction between light N1 sleep and wakefulness. The number in the morning is a model output, not a measurement.
Orthosomnia — when tracking becomes the problem
In 2017, researchers at Rush University Medical Center published a paper in the Journal of Clinical Sleep Medicine that named a pattern clinicians had started seeing in their patients. They called it orthosomnia: sleep anxiety induced by sleep tracking.
The patients weren’t sleeping badly because of an underlying disorder. They were sleeping badly because tracking had turned sleep into something to optimize, and optimizing sleep is almost perfectly designed to prevent it. Sleep is a biological process that improves when you stop trying and deteriorates when you try harder. The vigilant, evaluating mindset that makes optimization work in other domains is, at the physiological level, the mindset of an alert and active nervous system — which is to say, the opposite of what the transition into sleep requires.
Orthosomnia presents in a recognizable pattern. The person starts extending time in bed to increase their sleep-opportunity window (which often fragments sleep rather than deepening it). They start avoiding social commitments that would keep them up past a self-imposed bedtime. They check the score before they’ve assessed how they feel. They correlate a poor score with the next day’s performance, creating an anxious connection that tends to be self-fulfilling. Sleep anxiety that didn’t exist before the tracker started tracking is often orthosomnia.
This isn’t an argument against all tracking. It’s an argument for noticing whether the act of measuring is raising or lowering your background anxiety about sleep. If the number has become one more reason to lie in bed evaluating how well you’re doing it, the tracker is producing a net cost regardless of what it says.
The one metric worth watching
If there’s a single number a consumer sleep tracker reliably gives you that’s worth acting on, it is the consistency of your sleep timing — specifically, the variance in your wake time from day to day.
Not your deep sleep percentage. Not your sleep score. Not your REM duration, which varies normally across nights and in ways a wrist device can’t precisely distinguish from light sleep anyway. The time you wake up, and how much that time varies across the week.
Your circadian rhythm is a biological oscillator set by a cluster of neurons in the hypothalamus called the suprachiasmatic nucleus. Like any oscillator, it functions best when the inputs that reset it — primarily light, but also meal timing and physical activity — arrive at consistent intervals. Research on social jetlag, the mismatch between weekday and weekend sleep timing common in adults who use an alarm on workdays and sleep in on weekends, shows consistent associations with poorer mood, metabolic markers, and daytime cognitive performance, independent of total sleep duration.
In plain terms: sleeping eight hours a night but shifting your schedule by two hours on weekends produces something functionally like mild, weekly jetlag. The exhaustion you feel on Monday morning often isn’t from sleeping too little — it’s from sleeping at the wrong time relative to where your clock was set.
Consumer trackers are reasonably reliable at measuring sleep timing because timing is based on behavior — when you stopped moving, when you started again — rather than on inferences about brain states. A week’s graph of your wake times is accurate in a way your sleep stage breakdown probably isn’t.
The practice of pausing before checking the app isn’t about refusing information. It’s about giving your own felt sense of the night equal standing with the algorithm’s estimate, since your felt sense is often the more accurate of the two.
What to do when the score bothers you
A poor sleep score is not nothing, but it’s also not straightforwardly informative. Before assuming a bad night, it’s worth running through the variables the algorithm can’t fully see.
Was the timing unusual? A later-than-normal bedtime, an early alarm, or a time zone shift will produce a compressed or fragmented-looking night on most trackers regardless of actual sleep quality. Was there a physiological reason — alcohol, which measurably suppresses REM and fragments the second half of the night; illness, which changes HRV in ways algorithms weren’t trained on; a room that was too warm?
Was your felt experience consistent with the score? If you slept through and woke rested, and the app says 58, the mismatch is data about the algorithm, not about your night. Track the felt experience in parallel with the number over a few weeks. If they frequently diverge, weight the felt experience more. If they tend to agree, you have a useful signal.
The score is most meaningful when your sleep hygiene is already stable. When your sleep hygiene is inconsistent — irregular timing, variable pre-bed routine, light and temperature varying night to night — the score reflects those inconsistencies but can’t tell you which variable is driving it.
When to keep tracking and when to stop
There are genuinely useful reasons to track sleep for a defined period. If you’re trying to establish a consistent schedule, a week of data showing your timing variance is a clear picture of where to start. If you’ve made a behavioral change — shifting your wind-down by an hour, removing a late-evening habit — a two-week before-and-after comparison is real signal.
There are also signs that continuing isn’t helping. If you’re adjusting your schedule primarily to move a number rather than because the behavioral change itself seems worth making. If checking the app has become the first thing you do in the morning, before assessing how you actually feel. If a bad score on a good morning consistently undermines your sense of how the night went.
A 30-day snapshot is usually sufficient to identify the patterns that matter. Most of what there is to learn about your sleep timing, your weekday-to-weekend drift, and your average sleep window is visible within a month. Continuing to track beyond that is useful only if you have a specific question you’re using the data to answer.
The smallest version of this practice
Choose one consistent wake time for the next two weeks — not an aspirational earlier one, the one you can actually hold seven days a week. Set it. Don’t move it on weekends. Sleep onset tends to self-adjust over a couple of weeks as your circadian clock consolidates to the new anchor. If you use a tracker, use it to check one thing: whether your wake time is actually consistent. That question, asked once a week, is the most useful thing consumer tracking data can tell you.
The hours before that wake time are where the rest of the practice lives. What you take into the transition into sleep — the last conversation, the last screen, the last thought — shapes the night the tracker is trying to measure afterward. That’s where your bedtime routine and the sleep affirmations practice operate. Murmora’s format is built around the same window: a guide voice carrying something specific about your situation into the last minutes before sleep, quiet enough not to keep you awake and present enough to land. The tracker tells you something about what happened. This part shapes what happens.