# Wearables, PSG, and limitations of sleep measurement

Technology and future

Wearables often detect sleep better than quiet wakefulness

Actigraphy and optical pulse measurement provide indirect information about sleep. The analysis addresses accuracy compared to PSG, systematic errors, signal noise, and clinical application limits. Device version and the studied population are part of every accuracy specification.

## Validity is multidimensional

Chinoy and colleagues compared seven consumer devices in 34 healthy young adults with PSG. Sleep sensitivity was high, wake specificity significantly lower; stage results were inconsistent. [1] This structure explains why quiet wakefulness can appear as sleep. A high overall accuracy can also arise from the large proportion of actual sleep epochs. Sensitivity, specificity, temporal agreement, and total duration errors must be stated separately. A result from older devices is not a current test of every software today.

## Signal-to-noise ratio

Actigraphy records acceleration, not neural sleep activity. Movement can occur during sleep and wakefulness can be low in movement. Optical pulse measurement uses pulse-dependent light changes in the tissue. Motion artefacts, contact, blood flow, and light path influence the signal. Filters can reduce disturbances but also smooth out relevant changes. A clean pulse signal does not yet provide an EEG. Missing signal segments should not be hidden by an apparently precise sleep curve.

## Clinical limits and meaningful use

Another clinical comparison found limited agreement of apps and wearables with PSG. [2] A consumer watch can support trends or regular schedules without reliably diagnosing sleep apnoea, periodic movements, or psychiatric illness. Clinically approved individual functions must be assessed within their respective intended purpose. The product name alone does not allow equating with a complete PSG. In the case of complaints, the professional assessment counts, even if a watch shows good values. Conversely, a single bad grade is not proof of illness.

## An illustrative misclassification calculation

If a person sleeps for seven hours in an eight-hour recording, a device that labels every minute as sleep can already achieve 87.5 percent overall agreement. Yet, it does not detect a single minute of wakefulness. This pure calculation example explains why accuracy without class distribution is misleading. Confusion matrices and appropriate metrics are therefore used for stages. For total durations, bias and limits of agreement are important. A high correlation only means that values can vary together; two devices can systematically differ by a significant amount in doing so.

## From raw signal to decision

Filtering, feature extraction, and a classification model lie between the sensor and the visible sleep score. Every stage can lose information or introduce errors. A high correlation between two nocturnal mean values does not prove good agreement of individual events. Therefore, systematic deviation, scatter, and temporally matching comparison data are required for validation. The evaluation distinguishes between persons, nights, and short measurement epochs. Missing data must remain visible. A calm but awake person, poor skin contact, movement, or the signal of a partner can lead to plausible-looking misclassifications. For learning procedures, training and testing are separated by person. Otherwise, a model can recognise characteristics of the same person and overestimate its transferability. Devices and software versions are documented because an update can change the statement of an earlier validation.

## Feedback without self-reinforcing miscontrol

Automatic adjustment requires more than a good sensor. It must be determined when a change is sensible, how large it may be, and when the system should suspend a decision. Filters and delays reduce noise but may obscure rapid changes. Frequent corrections can themselves generate noise, movement, or alertness. A good control loop therefore limits actuation speed and interventions and offers an understandable manual return to a stable state. For an internal trial, sensor errors, actuator movements, and sleep events are logged on the same timeline. The comparison includes an unmodified setup and an actively controlled setup with the most similar expectation possible. An increase in the automatically calculated sleep score is insufficient if the same algorithm drives the regulation and evaluates its success. Independent endpoints and a traceable data concept are necessary. Only the signals required for the purpose are stored; a product benefit does not presuppose the permanent disclosure of all raw data.

Actigraphy | Movement-related sleep estimation | Quiet wakefulness problematic
PPG and movement | Additional physiological features | Artefacts and indirect stage inference
PSG | Multiple direct physiological channels | Effort and possible laboratory effects

Sensitivity | Proportion of correctly identified sleep epochs | High value can mask wakefulness errors
Specificity | Proportion of correctly identified wake epochs | Quiet wakefulness makes detection difficult
Duration deviation | Bias and limits of agreement | Correlation alone insufficient

A dedicated validation attempt synchronises raw signals and PSG, separates training and test subjects, and reports misclassifications per stage. Bland-Altman analysis complements correlations for total durations. People with insomnia, apnoea, and advanced age are specifically included. Missing data and software versions are disclosed.

STOLL can offer sensors as feedback with limitations. Changes to the mattress are not evaluated exclusively according to the sleep score of the same app. In the case of conspicuous symptoms, an unremarkable tracker should not prevent medical clarification.

## Actigraphy

Movement-related sleep estimation

Quiet wakefulness problematic

## PPG and movement

Additional physiological features

Artefacts and indirect stage inference

## PSG

Multiple direct physiological channels

Effort and possible laboratory effects

[1] Chinoy et al Performance of seven consumer sleep tracking devices compared with PSG
https://pubmed.ncbi.nlm.nih.gov/33378539/
34 healthy young adults and specific device versions; results do not apply generally to current models.

[2] Consumer grade sleep trackers compared with polysomnography
https://pubmed.ncbi.nlm.nih.gov/34741243/
Primary study; population and measurement methods limit the transferability to concrete bed products.

This paper is a targeted narrative research as of 30 September 2026. The starting point is the specific topic question, scientific publications, and, for technical or legal questions, the relevant original sources. The Word documents provided by the client serve as templates for the professional structure and comparative presentation. Their individual statements have not been adopted without verification. This research is not a systematic comprehensive survey, a meta-analysis, or a product certification.

The sources were checked via accessible publication sites, bibliographic datasets, and available excerpts. A complete article was not accessible for every source. Where only an abstract or excerpt was available, the description is limited to the information discernible therein. Figures are only mentioned within their study context; missing details are not supplemented. A phrase such as "no reliable evidence identified" describes the result of this targeted research and does not prove that no such work exists worldwide.

The source numbers in the text refer to the list at the end. Directly examined findings, mechanistic considerations, and the author's own practical deductions are linguistically separated. Hypothetical cases illustrate the decision-making logic; they are not documented customer experiences. The suggested test plans are original designs. They do not establish a binding standard or a medical treatment process. Statements about a product class are not automatically transferred to individual models.

For classification, the primary criterion is whether the source examines the exact question asked. A technically precise material measurement can be highly informative for a material property while saying little about sleep or long-term health. A clinical study may show a relevant benefit, but only for the group of people, construction, and duration of use studied. Proximity to the concrete question is therefore just as important as the study design.

Subsequently, comparison conditions, sample size, observation duration, and potential biases are considered. Blinding is often difficult with bedding. Expectations, habituation, and the sequence of tested variants can influence results. In the case of manufacturer funding, transparency and independent replication are particularly helpful; funding alone does not decide for or against the validity of a finding. Small pilot studies are primarily used to formulate a question more precisely and to plan a larger trial.

Statistical significance is not the same as practical importance. A small difference can be mathematically detectable without having a tangible benefit for the person in question. Conversely, a relevant individual improvement may remain statistically uncertain in a small group. Therefore, effect size, uncertainty, and everyday relevant endpoints are assessed together. A blanket score would obscure these differences. The interactive companion page consequently does not use fabricated health scores or simulated figures that appear like measured material data.

For implementation, a concrete goal is first defined, and then the smallest reasonably testable change is selected. The initial state, construction used, and observation period are documented. Feedback should capture both the desired benefit and possible new disadvantages. If several components are changed simultaneously, the attribution of success remains uncertain. An individual comparison can improve personal selection but does not replace a general efficacy study.

A supplier proof should concern the model actually offered and the intended use. Deviations in the cover, topper, base, care, or software can alter the transferability. The consultation openly states such limitations and formulates only the performance covered by data or immediate observation. For medical or legal questions, the relevant professional assessment remains necessary. The practical recommendation of this document is a basis for decision-making and not an individual diagnosis.