The Evidence
16,625,198 rated sessions from 868,244 people, collected over seven and a half years of ordinary use—every one rating their distress immediately before and immediately after. To our knowledge, the largest acute within-session characterization of a brief self-administered intervention reported to date.
16.6M
Rated sessions in the analytic cohort
868,244
Unique users contributing paired ratings
11 / 11
Distress domains showing a large effect
1.21–1.66
Range of effect sizes (Cohen's d)
Digital mental health interventions typically report effect sizes of d = 0.3–0.5 in clinical trials. Every one of the eleven domains here exceeds d = 1.2. The likely reason is when the intervention is reached for: this is not a controlled study run at convenient hours—it is what happens when people use something at the moment of actual need. Sleep sessions make the point plainly: the overwhelming majority are started late at night, not during research hours.
Cohen's d across digital mental health interventions
Effect size (Cohen's d): 0.2 = small, 0.5 = medium, 0.8 = large. Tapping figures are within-session change from this real-world analysis; comparison figures are from published controlled trials. The two are measured differently and are shown for scale.
Source: Ortner N, Nyer MB, Mischoulon D, Copeland DI, Fava M. Transdiagnostic Within-Session Outcomes of App-Delivered Tapping (Emotional Freedom Techniques): Real-World Evidence from 16.6 Million Rated Sessions. Depression Clinical and Research Program, Massachusetts General Hospital · Harvard Medical School · MIT Device Realization Lab. Manuscript under review—not yet peer-reviewed or published. Sterling IRB 16917-NOrtner. Reported following STROBE, RECORD, and STaRT-RWE standards. Data: October 2018 – 18 May 2026.
"EFT is the single most effective tool I've learned in 40 years of being a therapist."— Dr. Curtis Steele, Psychiatrist
Users rate their distress on a 0–10 scale before and after each session. This simple, ecologically valid measure captures the immediate impact of the intervention at the moment of use.
840 session titles were mapped to eleven clusters—among them anxiety, sleep, rumination, craving, and pain—by a classifier fixed before any effect was computed. Every one of the eleven showed a large within-session reduction, and the result held under six pre-specified sensitivity analyses, including restricting to each person's first session only, adjusting for regression to the mean, and recoding every missing post-rating as no change.
A second signal: response tends to be consistent within people across unrelated domains—someone who responds to tapping for anxiety tends to respond for pain and for craving too, suggesting responsiveness may be a property of the person rather than the presenting symptom.
The per-domain table with confidence intervals, the complete sensitivity analyses, and the methods are in the manuscript now under peer review. We will publish them here, and link the paper, once it has been through review. Evaluating NeuroTap for a population or a practice and need the detail sooner? Get in touch—we share the full analysis directly with clinical and payer partners under review.
Independent of Our Data
Everything above is our own real-world dataset. This is the separate, independent literature—other researchers, other cohorts, peer-reviewed and published. It is cataloged in full, nulls and criticisms included, at evidence.thetappingsolution.com: 407 studies and reviews, 99 randomized controlled trials, and 42 meta-analyses and systematic reviews across 41 countries.
A 2025 meta-analysis pooling ten randomized trials found significant reductions in both anxiety and depression against control. Zheng et al. (2025), Journal of Psychosomatic Research.
Hedges' g = 1.27 (95% CI 0.951–1.585, p < 0.001) across eighteen randomized trials. The authors note most participants had sub-threshold rather than major depression, and that two-thirds of included studies carried some risk of bias. Seok & Kim (2024), Journal of Clinical Medicine 13(21):6481.
A 147-patient randomized trial compared a six-week EFT program delivered online and self-paced against the same program delivered live. Pain severity, interference and quality of life improved against waitlist and held at six months—with no difference between self-paced and in-person delivery. Stapleton et al. (2025), European Journal of Pain 29(3):e4740.
Clond (2016), Journal of Nervous and Mental Disease: pre-post d = 1.23 against d = 0.41 for controls. Against other active treatments the difference (d = 0.44) did not reach significance—too few head-to-head studies to settle it. In the single direct comparison with CBT, the two performed comparably.
The pain trial is the one we would point a skeptical reviewer to first: randomized, six-month follow-up, and it tested the delivery model NeuroTap actually uses—a self-paced program on a screen, with no clinician in the room—against the same program in person. It found no advantage for the in-person version.
A result this size invites scrutiny. Here is exactly where the edges are, so you can judge it on the merits rather than take our word for anything.
The cohort is self-selected. There is no randomization, no control group, and no clinical diagnostic assessment. These data describe association, not causation.
Every measurement is taken immediately before and immediately after a single session. There is no follow-up beyond that session, so this evidence says nothing about durability over weeks or months.
The 0–10 scale parallels the numeric rating scale used for pain and the subjective units of distress used in behavior therapy, but it was not independently validated for every construct rated here.
Some session prompts used directional language. The sensitivity analysis suggests the contribution is small, but a residual effect cannot be ruled out.
This analysis has no depression cluster. The nearest constructs are Belief (d = 1.476—how true a limiting belief such as "I am not enough" feels) and Affect (d = 1.465—named distressing emotions). NeuroTap's depression evidence rests on the published trial literature, not on this paper.
The manuscript is under review. Nothing on this page has cleared peer review, and we will say so until it has.
What these data are good for: establishing that a brief, self-administered intervention produces a large, consistent, immediate change in rated distress across many domains, at a scale no trial could reach. What they are not: proof of efficacy. That requires the randomized controlled trials we are now building.
De-identified aggregate data supporting these findings are available to researchers on reasonable request, and we are actively seeking partners for the controlled trials.
Talk to Us About Research