Aller au contenu principal
Article language
STOLL/SleepBase/Technology and future

Technology and future

AI measurement and prediction of the suitable firmness level

A photo can support pre-selection but cannot determine a perfect mattress

Your gateway to sleep knowledge

Welcome to SleepBase.

Use your STOLL access to explore in-depth research on better sleep.

No access yet? Speak to your STOLL advisor.

Explore and understand

Discover connections.

Choose an approach, then explore each step with your mouse, keyboard or touch.

Photo

Potential contribution

Inexpensive geometric cues

Qualitative interpretation. These connections do not establish causality; no product measurements are simulated.

Change perspective

What does this mean in practice?

Only one photo

Geometry can be estimated.

Putting it in context

Mass, tissue, and preferences remain incomplete.

The research question

Understanding the findings.

Photos and body scans can capture geometric features. However, the suitable firmness level also depends on body mass, tissue distribution, sleeping position, physical complaints, preferences, and the specific mattress construction. An AI can only learn a recommendation as well as its target value and training data are defined. This analysis differentiates between body reconstruction, pressure prediction, and actual comfort prediction. For STOLL, a transparent assistance system with uncertainty specification and a subsequent lying test is more plausible than a seemingly exact diagnosis from a single image.

From image to geometry

A two-dimensional photo shows a projection. Perspective, clothing, lighting, and body posture influence the dimensions estimated from it. A three-dimensional scan can provide more shape information but can also have measurement errors. Neither a photo nor a surface scan automatically captures body mass or the mechanical properties of deep tissue.

The first quality question is, therefore, which features are actually measured and which are only estimated. Shoulder width, waist circumference, and visible contour are not the same as the load in the lateral position. An algorithm should not hide missing information behind an overly precise firmness specification.

Three different prediction tasks

Recognizing body posture, estimating pressure distribution, and predicting the most comfortable mattress are different tasks. The research paper BodyMAP investigates the joint estimation of a body model and three-dimensional pressure distribution for people in bed. [1] It shows technical possibilities but is not general proof of an ideal sales decision based on a normal customer photo.

Methods for sleep position recognition based on few pressure sensors also primarily address a classification task. [2] High accuracy for supine or lateral position says nothing about whether the recommended mattress reduces pain in the long term. A product must be evaluated according to exactly the goal it promises later.

The target value determines the result

If the model learns from previous sales, it can primarily reproduce previous sales practices. If it learns from pressure minima, it may optimize pressure at the expense of mobility or support. If it predicts satisfaction, brand, price, and service may influence the target variable.

A meaningful dataset should therefore contain several separate endpoints: comfort, position, movement, complaints, and, if applicable, objective measurements. Furthermore, the offered mattress range must be known. A firmness level H3 from a manufacturer is not a universal mechanical standard. The recommendation should be related to specific models or measured properties.

Validation outside of training data

A model must be tested on new individuals who do not appear in similar form in the training. If photos of the same person are distributed across training and testing, performance assessment can be too optimistic. New cameras, different clothing, or a changed product range can also reduce accuracy.

For STOLL, results based on body build, age, mobility, and relevant application situations would be important. High average accuracy can mask poor results for smaller groups. The system should be able to reject a recommendation or request additional measurements if the data lies outside its validated range.

Explainability and data protection

A useful explanation names the decisive features and the remaining uncertainty. "The shoulder zone might need more compliance" is more helpful for a lying test than an opaque overall rating. The final selection takes the person's feedback into account and remains correctable.

Body images are personal data and may contain particularly sensitive information depending on processing. Not every photo is automatically biometric data in the special legal sense; however, an identification or health analysis can change the requirements. Our own design recommendation is therefore to collect only necessary data, prefer local processing, and enable a consultation without a photo. [3]

Data leakage and seemingly high model accuracy

If multiple images of the same person appear in both training and testing, a model can use recognizable features. The test then appears more independent than it is. Similar problems arise if the subsequent target information is implicitly contained in the inputs already.

For fair evaluation, individuals and, if possible, recording conditions are cleanly separated. Additionally, an external test with new equipment or another location should take place. The accuracy must match the actual task: centimeter errors of a body measurement, classification accuracy of a position, and subsequent comfort are different results. A single high percentage cannot replace these differences.

Hypothetical case of an exact H3 recommendation

A photo algorithm recommends H3 with a seemingly 98 percent certainty. Without a definition of the target value, this number is hardly interpretable. Does it mean agreement with previous salesperson judgments, with a product category, or with long-term comfort?

STOLL should translate the recommendation into verifiable features. Which body region likely requires more compliance, which models are available for selection, and what uncertainty remains? A subsequent lying test can confirm or correct the hypothesis. An exact number is only helpful if its meaning and calibration are proven.

Recommendation system as a learning service

Feedback after purchase can improve a system, provided it is collected voluntarily, structurally, and in compliance with data protection. Complaints and cancellations should be considered just as much as satisfied customers. Otherwise, the model learns from a biased selection.

Changes to the product range, covers, or manufacturing processes can change the assignment. A model therefore requires versioning and re-testing. For STOLL, a transparently maintained assistance system is more plausible than a supposedly universal body formula trained once.

Benefits and limitations

Approaches in comparison.

ApproachPotential contributionLimits of the evidence
PhotoInexpensive geometric cuesPerspective and clothing limit analysis
3D scanMore detailed surfaceNo complete tissue mechanics
Pressure measurementLoad on a specific baseNo sole comfort standard
Lying test with feedbackCaptures personal fitShort-term and expectation-dependent

What can be measured.

MetricTestImportant limitation
Measurement errors of body dimensionsMeasure against referenceReport by body groups
Recommendation qualityNew individuals and productsNo data overlap
UncertaintyCalibrated specification or rejectionNo false precision
Long-term benefitFollow-up and complaintsSales success is not a health goal

From research to application

Guidance for practice.

A reliable test plan

A dedicated pilot should first compete against a simple, transparent consultation rule. The model receives only necessary features and recommends several plausible options instead of one allegedly perfect result. An independent group tests the suggestions without knowledge of the model's judgment. After several weeks, comfort and relevant complaints are recorded. Training and test individuals remain strictly separated. Changes to the product range trigger a re-evaluation. Implementation only occurs when the benefit compared to previous consultation and errors in important subgroups are comprehensible.

Implications for customer advice

STOLL can use AI for structured pre-selection, but should visibly differentiate between measurement, estimation, and recommendation. An alternative consultation without a body image remains useful. The customer receives a comprehensible justification and tries out the proposed systems. Statements about an ideal firmness level securely determined from a photo are not justified without extensive external validation.

Evidence and practical implementation

For classification, the primary criterion is whether the source examines the exact question asked. A technically precise material measurement can be highly informative for a material property while saying little about sleep or long-term health. A clinical study may show a relevant benefit, but only for the group of people, construction, and duration of use studied. Proximity to the concrete question is therefore just as important as the study design.

Subsequently, comparison conditions, sample size, observation duration, and potential biases are considered. Blinding is often difficult with bedding. Expectations, habituation, and the sequence of tested variants can influence results. In the case of manufacturer funding, transparency and independent replication are particularly helpful; funding alone does not decide for or against the validity of a finding. Small pilot studies are primarily used to formulate a question more precisely and to plan a larger trial.

Statistical significance is not the same as practical importance. A small difference can be mathematically detectable without having a tangible benefit for the person in question. Conversely, a relevant individual improvement may remain statistically uncertain in a small group. Therefore, effect size, uncertainty, and everyday relevant endpoints are assessed together. A blanket score would obscure these differences. The interactive companion page consequently does not use fabricated health scores or simulated figures that appear like measured material data.

For implementation, a concrete goal is first defined, and then the smallest reasonably testable change is selected. The initial state, construction used, and observation period are documented. Feedback should capture both the desired benefit and possible new disadvantages. If several components are changed simultaneously, the attribution of success remains uncertain. An individual comparison can improve personal selection but does not replace a general efficacy study.

A supplier proof should concern the model actually offered and the intended use. Deviations in the cover, topper, base, care, or software can alter the transferability. The consultation openly states such limitations and formulates only the performance covered by data or immediate observation. For medical or legal questions, the relevant professional assessment remains necessary. The practical recommendation of this document is a basis for decision-making and not an individual diagnosis.

Research you can trace

Sources and context.

Research methods and limitations

This paper is a targeted narrative research as of 30 September 2026. The starting point is the specific topic question, scientific publications, and, for technical or legal questions, the relevant original sources. The Word documents provided by the client serve as templates for the professional structure and comparative presentation. Their individual statements have not been adopted without verification. This research is not a systematic comprehensive survey, a meta-analysis, or a product certification.

The sources were checked via accessible publication sites, bibliographic datasets, and available excerpts. A complete article was not accessible for every source. Where only an abstract or excerpt was available, the description is limited to the information discernible therein. Figures are only mentioned within their study context; missing details are not supplemented. A phrase such as "no reliable evidence identified" describes the result of this targeted research and does not prove that no such work exists worldwide.

The source numbers in the text refer to the list at the end. Directly examined findings, mechanistic considerations, and the author's own practical deductions are linguistically separated. Hypothetical cases illustrate the decision-making logic; they are not documented customer experiences. The suggested test plans are original designs. They do not establish a binding standard or a medical treatment process. Statements about a product class are not automatically transferred to individual models.

English translation of the German original. Bibliographic references and Word files remain in their original language.

SleepBase by STOLL · Research status: 30 September 2026
Scientific interpretation with sources and limitations. Not an individual medical diagnosis.