I was recently asked to put together some thoughts on passive data and how it is used in healthcare. I thought I'd share a slimmed down version here.
Wearable devices have changed the way we think about health data. A watch can now capture sleep, heart rate, physical activity, respiratory rate and, in some cases, proxies for stress or recovery. Smartphones add another layer to this, including movement, location, screen use, communication patterns and environmental context. In mental health research, these forms of passive data are often presented as a way to move beyond traditional questionnaires and clinic-based assessments.
We often call it passive monitoring.
The appeal is obvious. Mental health is dynamic. Symptoms fluctuate across days, weeks and contexts. Someone’s mood, anxiety, sleep or alcohol use may change long before they reach a crisis point or speak to a clinician. Passive sensing through smartphones and wearables offers the possibility of more continuous, lower-burden monitoring in daily life, and recent reviews describe its potential for objective, non-invasive mental health monitoring. In depression research, this promise has been explored through work on digital health tools for the passive monitoring of depression, remote measurement in major depressive disorder, and studies examining whether smartphone and wearable data can help predict depression symptom severity.
But passive data does not save lives by itself.
Data only becomes useful when it is interpreted well, acted on appropriately, and embedded within ethical, trusted and clinically meaningful information systems. Without those conditions, passive data can just as easily mislead, exclude, overreach or create new forms of harm.
Passive does not mean neutral
One of the most common mistakes in digital health is to treat passive data as objective simply because it is automatically collected. A step count feels more real than a questionnaire response. A sleep score feels more precise than asking someone how rested they feel. A heart rate trace appears less subjective than a mood rating. But passive data is not neutral.
It is shaped by the device, the algorithm, the wearer, how it is worn, the setting and the assumptions built into the research question. Different devices measure the same physiological signal in different ways. Algorithms are often proprietary, meaning researchers may not know exactly how a sleep stage, stress score or recovery metric has been generated. Often when you ask simple questions of these companies, and the algorithms they have developed, they refuse to engage.
Data quality varies by device type, sensor placement, battery life, missingness, skin tone, activity type, occupational context and whether someone can afford or tolerate wearing the device consistently.
This is not just a theoretical concern. A 2025 scoping review of passive sensing and machine learning for mental health monitoring found promising work across depression, anxiety and other conditions, but also reported important limitations. Most studies were small, many had short monitoring periods, only one of 42 studies included external validation, and only a minority reported anonymisation clearly. Similar challenges have been reported in work using mobile health data from smartphones and wearable devices to predict depression symptom severity.
In mental health research, this matters. If we are trying to infer distress, relapse risk, depression severity or recovery from wearable data, we need to be honest about the limits of those inferences. Reduced activity may reflect low mood. It may also reflect injury, caring responsibilities, shift work, weather, poverty, fatigue, operational demands or simply a preference for rest. Poor sleep may be associated with anxiety or depression, and wearable-derived sleep features have been examined in relation to relapse in major depressive disorder.
It may also reflect parenting, pain, medication, noise, housing conditions or work patterns.
Passive data can tell us that something has changed. It cannot, on its own, tell us what that change means.
The signal is not always the construct
A major risk in wearable mental health research is that we confuse a physiological or behavioural signal with the psychological construct we care about.
Stress is a useful example. Many consumer wearables produce stress scores derived largely from physiological signals such as heart rate variability. These may be useful indicators of bodily arousal or recovery, but that is not the same as subjective stress, emotional distress or clinical deterioration.
Recent work comparing Garmin Vivosmart 4 data with repeated self-report measures in 781 students found robust associations for some sleep-related variables, weaker associations for tiredness, and little overlap between wearable and self-reported stress measures for most individuals. The authors concluded that wearable and self-report measures may not necessarily be measuring the same constructs. This distinction is also important in studies examining heart rate and sleep as potential markers of depression, including work on daytime and night-time heart rate dynamics as digital biomarkers of depression severity and nocturnal heart rate, sleep disruptions, anxiety and depression severity.
That finding should not make us dismiss wearables. Instead, it should make us more precise with our intent. We should not say a device measures stress when it measures physiological arousal. We should not say a sleep score measures recovery unless we understand how that score was produced and validated. We should not say a passive sensing model predicts depression unless we can explain what it predicts, for whom, under what conditions, and with what consequences.
This is where the ethics begin. Not only in the collection of data, but in the naming, interpretation and use of the signal.
The ethics sit in the interpretation
Much of the debate around wearable data focuses on consent and privacy. These are essential, but they are only part of the ethical picture. The more difficult questions often come later.
- What do we infer from the data?
- Who gets to see those inferences?
- What action follows?
- What happens if the inference is wrong?
In mental health, these questions are not abstract. A model might identify someone as being at increased risk of depression, alcohol misuse or psychological distress. But if that prediction is wrong, it could cause anxiety, stigma or inappropriate intervention. If the prediction is right but no support is available, the system may simply identify need without meeting it. If the data is used in an occupational setting, there may be concerns about surveillance, performance management or career consequences.
Ethics work in digital phenotyping has repeatedly highlighted privacy, data protection, consent, transparency, bias, fairness and accountability as central concerns for mental health applications. These are not administrative details. They determine whether a passive sensing system feels like support or surveillance.
This is particularly important in settings such as the Armed Forces, emergency services and other high-pressure workplaces. The potential value of passive data is clear. Earlier support, better understanding of occupational stressors, and more responsive health systems. But these are also settings where trust is fragile and where data can easily be perceived as something done to people rather than with them. Studies such as MAVERICK and V-MIND show the importance of carefully designed, population-specific approaches when using wearable and smartphone data with veterans.
An ethical wearable study is therefore not just one that has a consent form. It is one that has thought carefully about interpretation, governance, access, feedback, escalation, safeguards and participant trust.
From measurement to decisions
The phrase data saves lives is attractive. It also reflects a real policy direction, notably in the UK. The UK Government’s health and social care data strategy explicitly sets out a vision for making better use of data while reassuring people that their data will be handled safely and ethically.
But the phrase is incomplete. Data does not make decisions. People and systems do.
A wearable device might detect declining sleep, reduced activity and changes in heart rate variability. A smartphone might show reduced mobility or altered patterns of interaction. These signals may suggest that someone is struggling. But the important step is not the detection itself. It is the decision about what to do next.
- Should the person receive a prompt?
- Should they be asked to complete a brief check-in?
- Should a clinician be alerted?
- Should nothing happen unless the person has chosen that option in advance
- Should the system provide general support, signposting or personalised intervention?
Each of these decisions carries ethical weight. Too little action risks missing an opportunity to help. Too much action risks intrusion, false alarms and loss of autonomy. The right answer will not be the same for every person, population or context. This is why digital intervention research needs to focus not only on prediction, but also on the design and evaluation of response pathways, including work on personalised push notifications with AI, digital alcohol interventions for veterans, and the role of notifications in changing alcohol-related behaviours.
This is where digital health needs to mature. The goal should not simply be to collect more data or build more predictive models. The goal should be to design decision pathways that are proportionate, explainable, acceptable and useful.
Prediction is not the same as prevention
There is a tendency to assume that better prediction will automatically lead to better outcomes. Mental health is more complicated than that.
A systematic review of passive sensing for suicidal thoughts and behaviours found that passive sensing may be feasible in high-risk populations, but evidence for predictive value remained limited. The review identified 11 prediction studies and reported generally lower model performance for passive data compared with active data, with major shortcomings in methodology and reporting.
This is a useful warning. Even in areas where earlier detection could be lifesaving, we need to avoid overstating what passive data can currently do. Prediction models can produce risk scores, alerts and classifications, but prevention depends on what happens next. The quality of support, the acceptability of the intervention, the timeliness of response, and the person’s trust in the system.
A model that predicts risk but does not connect someone to meaningful support is not a safety net. It is an alarm without a response plan.
Consent needs to be ongoing
Passive data collection also challenges traditional models of consent. In many studies, and I am guilty of this in my own research, participants consent at the start, wear a device, and data collection continues in the background. That may be administratively efficient, but it does not necessarily reflect how people experience being monitored over time.
Someone may be comfortable sharing step count data but not location. They may agree to research use but not individual feedback. They may be happy for aggregated findings to inform policy but not for identifiable data to be used in employment-related decisions. Their preferences may also change as their circumstances change.
For passive sensing in mental health, consent should be treated as an ongoing relationship rather than a one-off transaction. Participants should understand what is being collected, what is not being collected, what can and cannot be inferred, who has access, and what happens if concerning patterns emerge. Longitudinal remote measurement studies such as RADAR-MDD and engagement-focused work such as RADAR-Engage show why retention, participant burden and ongoing engagement are not secondary details; they are central to whether these systems can work ethically and practically.
They should also have meaningful control where possible. Transparency is not simply about making information available. It is about making it understandable.
The risk of widening inequalities
Wearable research also risks widening health inequalities if we are not careful. People who own and regularly wear consumer devices may differ from those who do not. They may be healthier, wealthier, more digitally confident, or more comfortable with health tracking. Devices themselves may perform differently across populations. Some groups may be underrepresented in the datasets used to train models, meaning predictions may be less accurate for those who already experience poorer access to care.
There is evidence that wearable ownership is socially patterned. A large US survey of 23,974 respondents found that 44.5% owned a wearable device, with ownership higher among younger people, those with higher incomes, those with higher education levels, and those living in urban areas. The authors concluded that sociodemographic divides persist and that equitable access strategies are needed as wearables become part of clinical and public health domains.
There are also wider concerns about device performance and bias. The UK independent review on equity in medical devices found extensive evidence of poorer performance of pulse oximeters for patients with darker skin tones. Although pulse oximeters are not the same as consumer wearables, the lesson is directly relevant. Sensors, calibration datasets and validation studies can encode inequities if they are not designed and tested across diverse populations.
This matters because mental health research is already affected by inequalities in recruitment, diagnosis, treatment access and outcomes. If passive data systems are built primarily around those who are easiest to monitor, we may create tools that work best for the already visible.
Ethical wearable research therefore needs inclusive design from the start. That means involving participants, patients, clinicians, public contributors and communities in shaping the questions, methods and safeguards. It also means testing whether models work across different groups, rather than assuming that a signal identified in one population will generalise to another.
Better questions for wearable mental health research
The future of passive data in mental health should not be framed as a choice between optimism and scepticism. The technology has real potential. It may help identify early warning signs, personalise support, reduce recall bias, and improve our understanding of how mental health changes in daily life.
But we need to ask better questions.
Not just: can we predict depression?
But: should we predict it in this context, with these data, for this population, and with this pathway for support?
Not just: can we collect continuous data?
But: what level of monitoring is proportionate, acceptable and necessary?
Not just: does the model perform well?
But: for whom does it perform well, who might it fail, and what are the consequences of failure?
Not just: can data inform intervention?
But: who decides what intervention is appropriate, and how is the individual’s autonomy protected?
Not just: can the technology scale?
But: can it scale without eroding trust?
These questions are not barriers to innovation. They are what make innovation responsible.
Passive data is powerful, but not enough
Wearable and smartphone data will almost certainly play an increasing role in mental health research. Used well, these data can help us understand patterns that are difficult to capture through traditional methods alone. They can support more timely, personalised and context-aware approaches to care.
But passive data does not remove the need for judgement. It increases it.
The World Health Organization has argued that AI technologies for health must place ethics and human rights at the heart of design, deployment and use. The same principle should guide passive sensing in mental health. The question is not only whether we can collect more data, or whether models can detect patterns. The question is whether the systems we build are trustworthy, proportionate, inclusive and connected to meaningful support.
The ethical promise of wearable mental health research is not in the data itself. It is in the decisions made around it: what we collect, how we interpret it, who we involve, what we feed back, how we protect people, and whether the resulting systems genuinely improve care.
Passive data does not save lives. Ethical, trusted and evidence-based decisions made from it might.