Research / HealthTech
·8 min read
71% of healthtech startups use AI.How do they prove their software actually works?
AI is an integral part of the product at 20 of 28 healthtech startups with Belarusian roots: it personalises workouts, analyses nutrition, recognises skin lesions or heart rhythms. We looked at how companies test their technologies, from algorithm accuracy and user outcomes to clinical studies and medical certification.
Clinical testing
Technology validation
Assessing user outcomes
Academic mentions
Regulation
/ certification
- 20 / 28
- Use AI as part of the product
- 12 / 28
- Have a public scientific, clinical or regulatory footprint
- 16 / 28
- No public signals found
Contents
00About the research
This is the second article based on HealthTech in Motion, a study of international healthtech startups founded by Belarusians. In the first part, we explored who is building health ventures today and why three quarters of the companies are started by experienced founders. This time, we look at what stands behind their technologies. According to the research, no public scientific, clinical or regulatory signals were found for 16 of the 28 companies (57%). This does not mean their technologies are untested or do not work. The remaining 12 have a public footprint, and it takes different forms.
- Have a public scientific, clinical or regulatory signal
- 12
- No public signals found
- 16
The absence of a public signal does not mean a technology is untested or does not work.
01Technology validation
When the technology itself is tested
Skinive, for example, an AI service for identifying skin conditions, tested the accuracy of its pathology recognition algorithm. In a published paper, the company compared two versions trained on 64,000 and 115,000 images. For malignant lesions, the newer version reached a sensitivity of 95.4%.
HeartScan, an app for monitoring heart activity using a smartphone, tested how its algorithm detects cardiac cycle events. It used 7,700+ hours of accelerometer recordings collected from different smartphone models under noisy conditions and during movement. The best version of the algorithm detected the moment of aortic valve opening with 99.8% accuracy. The paper is currently available as a preprint.
Doctorina, an AI service for medical consultations, took a different approach: the team compared its algorithm’s diagnostic results with those of doctors and other AI models. It used 150 simulated primary care consultations. Doctorina’s diagnosis matched the reference diagnosis in 82% of cases, compared with 57% for doctors. The startup’s algorithm also achieved the highest diagnostic accuracy against frontier models (Claude Opus 5, GPT-5.6, Gemini 3.1 Pro, Kimi K3). This study is also currently available as a preprint.
Skinive
95.4%
HeartScan
99.8%
Doctorina
82% vs 57%
These are different measures on different datasets. The percentages cannot be compared as a product ranking.
02Assessing user outcomes
When research goes beyond the algorithm
Flo Health, a women’s health app, studies different aspects of its product. Two pilot randomised studies involving more than 400 women, for example, examined the role of educational content in menstrual health, PMS and PMDD. After the three-month studies, participants had a better understanding of women’s health and reported improvements in some indicators. The group with PMS and PMDD also reported less severe symptoms. The company also tested the accuracy of its Symptom Checker using simulated cases of endometriosis, fibroids and polycystic ovary syndrome. Its conclusions agreed with assessments by independent doctors in 83–88% of cases. This year, Flo registered a study of its digital contraception feature.
Goodville, a mobile game that uses game mechanics to support mental health, studied changes in users rather than an algorithm. The company conducted a six-week study involving more than 1,700 people with symptoms of depression. Around 60% reported a reduction in symptom severity during that period. The study was observational, however, with no randomisation or control group. It recorded changes in users over that time, but did not prove that the app caused them.
CleverPoint, a platform for assessing psychophysiological state using EEG/ECG and virtual reality, served as a research tool rather than the subject of the research. It was used to assess how the brain and body responded to cognitive load in 400+ ship officers and engineers. Significant stress was detected in 73.8% of participants, while participants over 40 showed a decline in cognitive function under stress. The study therefore examined people’s condition using the technology, rather than CleverPoint’s accuracy.
Flo Health
- Women in two pilot randomised studies
- 400+
- Agreement between Symptom Checker and independent doctors’ assessments
- 83–88%
Digital contraception study registered
Goodville
~60%
CleverPoint
73.8%
03Clinical testing
From a pilot with 40 children to testing on 1,000 patients
According to the research, three more startups are at different stages of clinical testing. Bloom, an AI nutritionist and food tracker, mentions testing by clinics in its investor materials. VR NeuroLink AI, a developer of VR technology for neurorehabilitation, reports testing its product on 1,000+ patients. MindMuscle, a neurofeedback platform and EEG headset for children with ADHD, reports a pilot with 40+ children and work with five paediatric clinics in the US.
Since the methodology and results of these tests are not publicly available, it is impossible to determine what they showed. These cases are therefore counted separately from published scientific papers in the research.
Bloom
Clinics
VR NeuroLink AI
1,000+
MindMuscle
40+
Methodology and results are not publicly available. These cases are counted separately from published scientific papers.
04Academic footprint
When researchers take notice of a startup
Several other startups in the sample have appeared in independent scientific papers. A presence in academic literature does not, however, establish a product’s effectiveness: a review may contain a full analysis or simply mention it as an example.
Skinive, for instance, has attracted independent research attention at least twice: the app was compared with other AI skin analysis services in terms of stated accuracy, doctors’ involvement, regulatory status and how fully it discloses information to users. Stork, a pregnancy tracking app, was also analysed alongside other products in an independent study of pregnancy self-management apps.
Soula, meanwhile, was cited as an example of an existing solution, and data from Zing Coach’s research was used in an academic publication without an assessment of the app’s own effectiveness.
- Skinive
- Comparison with other AI skin analysis services
- Stork
- Analysis in a study of pregnancy self-management apps
- Soula
- Mention as an example of an existing solution
- Zing Coach
- Use of research data without evaluating the app’s effectiveness
Appearing in a scientific paper does not by itself confirm a product’s effectiveness.
05Regulation, ISO, patents
Other signs of maturity: regulation, ISO, patents
For some healthtech products, research is only part of the process: if a startup intends its product to be used for medical purposes, it will have to go through the medical device regulatory process. Two companies in the sample are on this path: Skinive has CE marking as a class I medical device under the EU MDR, and VR NeuroLink AI is undergoing certification as a class IIa medical device.
Other signs of company maturity include patents and management standards. Skinive is certified to ISO 13485, a quality management standard for medical device development. Flo Health holds ISO 27001 and ISO 27701 certifications for information security and data privacy. According to the company, it is the only femtech app certified to both standards. MindMuscle has obtained a US patent for its EEG headset.
- Skinive
- CE · class I
- Medical device under the EU MDR. ISO 13485: quality management for medical device development.
- VR NeuroLink AI
- Class IIa · in progress
- Undergoing medical device certification.
- Flo Health
- ISO 27001 · ISO 27701
- Information security and data privacy.
- MindMuscle
- US patent
- For its EEG headset.
06Expert comment
When good metrics are no longer enough
As these examples show, there is no single ladder of evidence in HealthTech. An algorithm that analyses heart signals and medical software that assesses skin lesions perform different tasks. They therefore require different studies and levels of evidence. Where is the boundary, and when does a digital health product need clinical validation at all? We asked the head of a technology company with many years of experience developing digital health products for the US market, including projects for major hospitals and healthcare enterprises. The company has delivered several hundred projects. The expert asked to remain anonymous.
Retention only tells you that people use the product. It does not prove the product’s effectiveness.
The expert asked to remain anonymous
— At what point is it no longer enough for a healthtech startup that users simply like the product and its metrics look good? What usually makes a company move towards clinical validation and begin the regulatory process?
— In short, in our experience, metrics are not what push a startup towards clinical validation. Two factors are decisive: the startup’s wish to make medical claims about its product, and the arrival of a payer other than the user, such as a clinic, insurer or employer. Retention only tells you that people use the product. It does not prove the product’s effectiveness, which is precisely what the new buyer wants to know.
— How much can stated functionality, or claims, change a product? For example, one startup helps people better understand their sleep, while another identifies the risk of a disease. Technically, they may use similar AI models. What fundamentally changes in the second case?
— Claims, the promises a startup makes to users and the market, define the product’s intended use. That intended use determines whether the product is considered a medical device. The model itself may be identical, while the products are different. Depending on the claims, everything around the model changes: the cost of an error, data requirements, the development process and even the freedom to update the model.
— Startups often begin as consumer, wellness or B2C products, then want to work with clinics, insurers and doctors and encounter a completely different level of evidence. What usually needs to be rebuilt? And what should they plan for from the start?
— Most of the rework can be avoided, and doing so can be relatively inexpensive. Study the industry before writing code, and keep the path into medicine open even if you are building a wellness product today. If you overlook this, you will later have to rebuild the foundations rather than the interface: data, architecture, the evidence base and the business model.
— What is the most expensive or difficult part of moving from a functioning healthtech product to one that must withstand clinical and regulatory scrutiny: data, research, product architecture, documentation, processes or something else?
— Honestly, and with a little irony, the most expensive thing is the time already spent on what now has to be redone. More seriously, the biggest financial costs are associated with clinical evidence. In terms of risk, the most dangerous thing is choosing the wrong intended use or trying to retrofit requirements to a product originally developed without them.
For more on what startups should consider before their product needs clinical validation — data, architecture, documentation and the business model — look out for an upcoming article on BYGRID.