16.6M

Rated sessions in the analytic cohort

868,244

Unique users contributing paired ratings

11 / 11

Distress domains showing a large effect

1.21–1.66

Range of effect sizes (Cohen's d)

Why these numbers matter

Digital mental health interventions typically report effect sizes of d = 0.3–0.5 in clinical trials. Every one of the eleven domains here exceeds d = 1.2. The likely reason is when the intervention is reached for: this is not a controlled study run at convenient hours—it is what happens when people use something at the moment of actual need. Sleep sessions make the point plainly: the overwhelming majority are started late at night, not during research hours.

Effect Size Comparison

Cohen's d across digital mental health interventions

Tapping — Craving
d = 1.66
Tapping — Sleep
d = 1.49
Tapping — Anxiety
d = 1.44
Tapping — Pain
d = 1.21
Daylight (Big Health)
d = 1.08
Typical Digital CBT
d = 0.3–0.5
Meditation Apps
d = 0.2–0.4

Effect size (Cohen's d): 0.2 = small, 0.5 = medium, 0.8 = large. Tapping figures are within-session change from this real-world analysis; comparison figures are from published controlled trials. The two are measured differently and are shown for scale.

Source: Ortner N, Nyer MB, Mischoulon D, Copeland DI, Fava M. Transdiagnostic Within-Session Outcomes of App-Delivered Tapping (Emotional Freedom Techniques): Real-World Evidence from 16.6 Million Rated Sessions. Depression Clinical and Research Program, Massachusetts General Hospital · Harvard Medical School · MIT Device Realization Lab. Manuscript under review—not yet peer-reviewed or published. Sterling IRB 16917-NOrtner. Reported following STROBE, RECORD, and STaRT-RWE standards. Data: October 2018 – 18 May 2026.

"EFT is the single most effective tool I've learned in 40 years of being a therapist."
— Dr. Curtis Steele, Psychiatrist

How we measure outcomes

Users rate their distress on a 0–10 scale before and after each session. This simple, ecologically valid measure captures the immediate impact of the intervention at the moment of use.

840 session titles were mapped to eleven clusters—among them anxiety, sleep, rumination, craving, and pain—by a classifier fixed before any effect was computed. Every one of the eleven showed a large within-session reduction, and the result held under six pre-specified sensitivity analyses, including restricting to each person's first session only, adjusting for regression to the mean, and recoding every missing post-rating as no change.

A second signal: response tends to be consistent within people across unrelated domains—someone who responds to tapping for anxiety tends to respond for pain and for craving too, suggesting responsiveness may be a property of the person rather than the presenting symptom.

Full results are held for publication

The per-domain table with confidence intervals, the complete sensitivity analyses, and the methods are in the manuscript now under peer review. We will publish them here, and link the paper, once it has been through review. Evaluating NeuroTap for a population or a practice and need the detail sooner? Get in touch—we share the full analysis directly with clinical and payer partners under review.

The published, peer-reviewed evidence

Everything above is our own real-world dataset. This is the separate, independent literature—other researchers, other cohorts, peer-reviewed and published. It is cataloged in full, nulls and criticisms included, at evidence.thetappingsolution.com: 407 studies and reviews, 99 randomized controlled trials, and 42 meta-analyses and systematic reviews across 41 countries.

Anxiety — ten RCTs, 774 patients

A 2025 meta-analysis pooling ten randomized trials found significant reductions in both anxiety and depression against control. Zheng et al. (2025), Journal of Psychosomatic Research.

Depression — eighteen RCTs

Hedges' g = 1.27 (95% CI 0.951–1.585, p < 0.001) across eighteen randomized trials. The authors note most participants had sub-threshold rather than major depression, and that two-thirds of included studies carried some risk of bias. Seok & Kim (2024), Journal of Clinical Medicine 13(21):6481.

Chronic pain — self-paced matched in-person

A 147-patient randomized trial compared a six-week EFT program delivered online and self-paced against the same program delivered live. Pain severity, interference and quality of life improved against waitlist and held at six months—with no difference between self-paced and in-person delivery. Stapleton et al. (2025), European Journal of Pain 29(3):e4740.

Anxiety — the most-cited meta-analysis

Clond (2016), Journal of Nervous and Mental Disease: pre-post d = 1.23 against d = 0.41 for controls. Against other active treatments the difference (d = 0.44) did not reach significance—too few head-to-head studies to settle it. In the single direct comparison with CBT, the two performed comparably.

The pain trial is the one we would point a skeptical reviewer to first: randomized, six-month follow-up, and it tested the delivery model NeuroTap actually uses—a self-paced program on a screen, with no clinician in the room—against the same program in person. It found no advantage for the in-person version.

What this data does not establish

A result this size invites scrutiny. Here is exactly where the edges are, so you can judge it on the merits rather than take our word for anything.

Observational, with no control arm

The cohort is self-selected. There is no randomization, no control group, and no clinical diagnostic assessment. These data describe association, not causation.

Within-session only

Every measurement is taken immediately before and immediately after a single session. There is no follow-up beyond that session, so this evidence says nothing about durability over weeks or months.

Self-reported, single-item ratings

The 0–10 scale parallels the numeric rating scale used for pain and the subjective units of distress used in behavior therapy, but it was not independently validated for every construct rated here.

Demand characteristics can't be fully excluded

Some session prompts used directional language. The sensitivity analysis suggests the contribution is small, but a residual effect cannot be ruled out.

Depression is not directly measured here

This analysis has no depression cluster. The nearest constructs are Belief (d = 1.476—how true a limiting belief such as "I am not enough" feels) and Affect (d = 1.465—named distressing emotions). NeuroTap's depression evidence rests on the published trial literature, not on this paper.

Not yet peer-reviewed

The manuscript is under review. Nothing on this page has cleared peer review, and we will say so until it has.

What these data are good for: establishing that a brief, self-administered intervention produces a large, consistent, immediate change in rated distress across many domains, at a scale no trial could reach. What they are not: proof of efficacy. That requires the randomized controlled trials we are now building.

Want the underlying data?

De-identified aggregate data supporting these findings are available to researchers on reasonable request, and we are actively seeking partners for the controlled trials.

Talk to Us About Research