Key Points 5 min read
  • A huge sample and large effect size prove nothing without a control group — that's correlation, not causation
  • Single-item self-ratings buy engagement and sample size, but trade away measurement rigor and objectivity
  • Pre/post designs invite regression to the mean and expectation bias, inflating effects that never happened
Table of Contents

If you saw a headline — “84,000 users, an app significantly reduces stress, effect size 0.7” — what should your first reaction be?

Not “great, downloading it,” and not “must be an ad.” It’s: look at the study design first. A December 2025 npj Digital Medicine paper on a self-hypnosis app is a perfect exercise for exactly this — the numbers are pretty, but underneath them sits a row of question marks that deserve to be raised.

TL;DR

  • Subject: the self-hypnosis app Reveri (Stanford’s David Spiegel is co-founder / scientific advisor) — 84,395 users, 282,893 stress sessions, Nov 2021–Jan 2025
  • Result: a single-item 10-point Likert self-rating before/after each session; across the first 10 sessions, Cohen’s d = −0.71 to −0.78 (statistically a “large” effect), mean drop 1.45–1.62 points
  • But it’s a retrospective observational study, no control group, self-reported single-item measure — which keeps it at “correlation,” and it can’t prove hypnosis caused the change
  • For engineers / anyone reading data: a classic case of “big N ≠ rigorous” and “big effect size ≠ causation”

What the study actually did

It’s not a randomized controlled trial (RCT) — it’s a retrospective observational study: take the real usage data the app accumulated and analyze it. Users self-rated stress before and after each session (1 = lowest, 10 = highest); the study looked at “after − before.”

The authors openly state why they used a single-item Likert instead of a multi-item validated questionnaire: to maximize engagement and data collection. An honest but crucial tradeoff — you trade a coarse measurement for an 84k sample.

The numbers themselves

Across the first 10 sessions, Cohen’s d sits at −0.71 to −0.78. By convention d≈0.2 is small, 0.5 medium, 0.8 large — so this is close to “large.” Mean stress dropped 1.45–1.62 points (out of 10).

Some interesting moderators:

  • Higher hypnotizability → bigger drop — but the correlation is actually weak (ρ ≈ −0.11 to −0.15)
  • Interactive / standard-length sessions were 1.8× as effective as the brief version (−2.02 vs −1.13)
  • Younger users dropped more on session 1; older users got more cumulative benefit
  • Paying members dropped more in sessions 1–10
  • Repeating sessions did NOT compound the effect over time

Safety looked good: of 84,395 users, only 10 reported worsening symptoms or other problems, all minor.

Why a pretty number still isn’t a conclusion

This is the point. Each item below is a reason d=0.7 loses its persuasive power:

1. No control group. The most fatal. You only know “users self-rate lower stress after a session” — you don’t know whether they’d have dropped just as much doing nothing, or sitting with eyes closed for 5 minutes. The authors themselves admit this limits causal inference.

2. Regression to the mean. People usually open a stress app when stress is high. An unusually high state tends to drift back toward average on the next measurement — even with no intervention. A pre/post design bakes this natural rebound into the “effect.”

3. Self-report + single item + expectation. The user knows they’re using a stress tool, knows the question is about stress, and just spent time on something that “should” help. How much of the drop is real physiological relaxation vs “I expected it to work, so I gave a lower number”? A one-item Likert can’t tell them apart.

4. Selection bias. The sample is “people who downloaded a paid self-hypnosis app and bothered to fill in pre/post ratings” — heavily self-selected, not representative. Paying members doing better is likely just “more invested people feel more effect.”

5. A big sample doesn’t fix any of the above. Many see N=84,395 and think “surely that many people is trustworthy.” But a large sample only makes your estimate more precise; it doesn’t make a biased design unbiased. Regression to the mean at 84k is still regression to the mean — you just measure a contaminated number with great confidence. The authors say the large sample “helps mitigate” confounding; that’s true for random error, not for systematic bias.

So is the study worthless?

No. Put it back where it belongs and it’s valuable:

  • It shows real-world usage data can be systematically collected and analyzed (the data backbone of digital therapeutics)
  • It gives a signal and an effect-size range worth confirming with a follow-up RCT
  • The safety data (10/84,395) is genuinely informative at that scale
  • The moderators (interactivity, hypnotizability) have real product-design implications

It just cannot be read as “the hypnosis app is proven to reduce stress.” It can be read as “among self-selected users, self-rated stress drops after a session — worth a controlled trial.”

What to take away (the data-literacy version)

Next time you see a “huge N, beautiful effect size” health/behavior study, ask in order:

  1. Is there a control group? No → probably only correlation.
  2. Is it pre/post? Yes → watch for regression to the mean and expectation effects.
  3. How was it measured? Self-reported single item → subjective, expectation-prone.
  4. Where did the sample come from? Self-selected → don’t extrapolate.
  5. Is the big sample solving random error, or being used to paper over a design flaw?

The same instinct applies in engineering: an A/B test without a control isn’t an A/B test. A pretty metric with no control and no randomized split usually shows you regression to the mean and selection effects — not the merit of your change. Reading papers and reading dashboards require the same caution.

References

Ask this article

Answers come from this article only. Click any prompt below or open the chat at the bottom right.

🇺🇸 English

If you saw a headline — "84,000 users, an app significantly cuts stress, effect size of 0.7" — what should your very first reaction be?

It shouldn't be "great, downloading it." And it shouldn't be "must be an ad." It should be: let me look at how the study was actually designed. Because there's a paper from December 2025, in npj Digital Medicine, about a self-hypnosis app — and it's a perfect workout for exactly that instinct. The numbers look gorgeous. But sitting right underneath them is a whole row of question marks, and those deserve some daylight.

So let's set the stage. The app is called Reveri — and one of its co-founders and scientific advisors is David Spiegel, from Stanford. The dataset is big: over 84,000 users, nearly 283,000 stress sessions, collected from late 2021 through the start of 2025. Here's how the measurement worked. Before each session and after each session, users rated their own stress on a single ten-point scale — one is lowest, ten is highest. The study looked at the difference: after, minus before.

And the authors are refreshingly honest about why they used a single one-item rating instead of a proper, validated, multi-question questionnaire. Their reason? To maximize engagement and data collection. That's a real tradeoff, and worth pausing on — they traded measurement precision for an eighty-four-thousand-person sample. Coarser ruler, way more data.

Now the headline numbers. Across the first ten sessions, the effect size — Cohen's d — lands between negative 0.71 and negative 0.78. For reference, the rule of thumb is: 0.2 is small, 0.5 is medium, 0.8 is large. So this is knocking on the door of "large." In plain terms, average stress dropped about one and a half points out of ten.

A few interesting wrinkles in the details. People who scored as more hypnotizable tended to drop a bit more — but honestly that link was weak. The interactive, full-length sessions worked noticeably better than the quick brief ones — roughly 1.8 times the effect. Younger users got a bigger hit on their very first session; older users built up more benefit over time. Paying members improved more. And here's a curious one — repeating sessions did not stack the effect. Doing it over and over didn't compound. On safety, it looked clean: out of all 84,000-plus users, only ten reported anything getting worse, and all of it minor.

Okay. Now the heart of the whole thing — why a pretty number still isn't a conclusion. Let me walk you through the reasons, because each one chips away at that 0.7.

Reason one, and it's the fatal one: there's no control group. None. All we actually know is that people rate their stress lower after a session. What we don't know is whether they'd have dropped just as much doing nothing at all — or just sitting quietly with their eyes closed for five minutes. The authors admit this straight up: it limits any claim about cause.

Reason two: regression to the mean. Think about when you actually open a stress app. You open it when you're stressed. That's a high, unusual moment — and unusually high states naturally drift back toward your average the next time you check, with zero intervention. A before-and-after design bakes that natural rebound right into what looks like the "effect."

Reason three: it's self-reported, it's a single item, and expectation is in the room. The user knows they're using a stress tool. They know the question is about stress. They just spent time on something that's supposed to help. So how much of that drop is genuine physiological calm, versus "I expected this to work, so I typed a lower number"? A one-question scale simply cannot separate those two things.

Reason four: selection bias. Who's in this sample? People who downloaded a paid self-hypnosis app and were diligent enough to fill out ratings before and after. That's a heavily self-selected crowd — not a slice of the general public. And remember paying members did better? That's probably just "more invested people feel more effect."

And reason five — this is the one I really want to land. A big sample does not fix any of the previous four. People see N equals 84,000 and think, surely that many people can't be wrong. But here's the thing: a large sample only makes your estimate more precise. It does not make a biased design unbiased. Regression to the mean at 84,000 people is still regression to the mean — you've just measured a contaminated number with tremendous confidence. The authors say the large sample helps mitigate confounding, and that's true for random noise — but not for systematic bias. Precision and accuracy are not the same animal.

So — is the study worthless? No. Not at all. Put it back where it belongs and it's genuinely valuable. It shows you can systematically collect and analyze real-world usage data — that's the backbone of digital therapeutics. It hands you a signal and an effect-size range that are absolutely worth confirming with a proper randomized trial. The safety data, ten out of 84,000, is actually informative at that scale. And the moderators — the fact that interactive sessions worked better — have real product design implications. What it can't do is be read as "the hypnosis app is proven to reduce stress." What it can be read as is: among self-selected users, self-rated stress drops after a session, and that's worth a controlled trial.

So let me leave you with the part you can actually carry around — the data-literacy checklist. Next time you see a "huge sample, beautiful effect size" health or behavior study, ask these in order. One: is there a control group? If no, you're probably looking at correlation, full stop. Two: is it before-and-after? If yes, watch for regression to the mean and expectation effects. Three: how was it measured? Self-reported single item means subjective and expectation-prone. Four: where did the sample come from? Self-selected means don't extrapolate. And five: is that big sample solving random error, or is it being used to paper over a design flaw?

And here's the kicker — the exact same instinct applies in engineering. An A/B test without a control isn't an A/B test. A gorgeous metric with no control and no randomized split is usually showing you regression to the mean and selection effects — not the actual merit of your change. Reading papers and reading dashboards take the same caution.

So, three things to walk away with. First: a big N and a big effect size prove neither rigor nor causation — without a control group, you've got correlation, dressed up nicely. Second: before-and-after measurement quietly invites regression to the mean and expectation bias, especially when people show up precisely because they feel bad. And third: a large sample buys you precision, never accuracy — it'll measure a flawed design with beautiful, misleading confidence. Look at the design first. Always the design first.

🇹🇼 中文

如果你在新聞上看到一句話:「八萬人實測,某 App 顯著減壓,效果量零點七」,你的第一反應該是什麼?

不是「太好了我要下載」,也不是「這一定是業配」。正確的反應是——先看研究怎麼設計的。

二零二五年十二月,npj Digital Medicine 登了一篇關於自我催眠 App 的論文,剛好就是練習這件事的好教材。數字很漂亮,但漂亮的數字底下,藏著一整排該打的問號。

先講這篇研究到底做了什麼。

它分析的是一個叫 Reveri 的自我催眠 App,背後的共同創辦人是史丹佛的 David Spiegel。資料量很驚人:八萬四千多名使用者、超過二十八萬次的減壓 session,時間橫跨二零二一年底到二零二五年初。

關鍵來了——它不是隨機對照試驗,而是回溯性的觀察研究。意思是,作者直接拿 App 累積下來的真實使用資料來分析,沒有實驗室、沒有分組。使用者在每次 session 的前後,各自評一次壓力,一分最低、十分最高,研究就看「做完之後比做之前降了多少」。

而且測壓力用的是單題的 Likert 量表,不是多題的標準心理量表。作者很誠實,直接講了理由:為了把參與率跟資料量衝到最高。這是個關鍵的取捨——你用最低摩擦的測量,換到了八萬人的規模,但代價就是這個測量本身很粗糙。

那數字長怎樣?

前十次 session,效果量 Cohen's d 落在負零點七一到負零點七八之間。按照慣例,零點二是小效果、零點五中等、零點八算大,所以這已經很接近「大」效果了。平均壓力下降一點五分左右,滿分十分。

裡面還有幾個蠻有意思的發現。可被催眠的程度越高,壓力降得越多——但其實這個相關性很弱。互動式、完整長度的 session,效果是簡短版的一點八倍。年輕人第一次降得多,年長者反而是累積效果比較好。付費會員前十次也降得比較多。

但有一個反直覺的點:重複做,並不會讓效果越疊越高,沒有所謂的長期累積效應。

安全性看起來很好,八萬多人裡只有十個人回報症狀變差,而且都很輕微。

好,現在進入這集真正的重點——為什麼數字這麼漂亮,我們還是不能下結論?下面每一條,都是讓「零點七」這個效果量失去說服力的理由。

第一,也是最致命的:沒有對照組。你只知道使用者做完 session 後,自評壓力降了。但你完全不知道——如果他們什麼都不做,或只是閉著眼睛坐五分鐘,會不會也降一樣多?沒有對照組,你就沒有一個「基準」可以比。作者自己也承認,這限制了因果推論。

第二,迴歸均值。這個概念很重要。人通常是在壓力特別高的時候,才會打開減壓 App,對吧?而「特別高」這種極端狀態,下一次再量,本來就傾向往平均值靠回去——就算你中間什麼都沒做。所以前後測這種設計,天生就會把這種自然回落,算進「效果」裡面。

第三,自陳測量加上期望效應。使用者很清楚自己正在用一個減壓工具,也知道這題在問壓力,而且剛花了時間,心裡覺得「應該要有效」。在這種情境下,自評分數降下來,有多少是真的生理放鬆,有多少只是「我期待它有效,所以我給個低分」?單題的 Likert,根本分不出來。

第四,選擇偏誤。這個樣本是哪些人?是會自己跑去下載、還願意付費用自我催眠 App,而且乖乖填前後測的人。這是高度的自我選擇,沒辦法代表一般大眾。付費會員效果比較好,很可能也只是因為——更投入的人,本來就更容易感受到效果。

第五,這條最反直覺:大樣本治不了上面這些問題。很多人一看到 N 等於八萬四千,就覺得「這麼多人,總該可信了吧」。但是,樣本大,只會讓你的估計更精準,不會讓一個有偏誤的設計變得無偏。八萬人的迴歸均值,它還是迴歸均值;你只是會非常有信心地,測到一個被污染的數字。作者說大樣本「有助於緩解」混淆,這句話對隨機誤差是成立的,但對系統性偏誤,不成立。

那講到這,你可能會問:所以這篇研究是廢的嗎?

不是。把它擺回它該在的位置,它其實很有價值。它證明了真實世界的使用資料,是可以被系統性地收集跟分析的,這是數位療法的資料基礎。它給了後續的隨機對照試驗,一個值得去驗證的訊號跟效果量範圍。安全性資料在這種規模下也很有參考價值。那些調節變項,像互動性、可催眠程度,對產品設計也有實際的啟發。

它只是不能被讀成「催眠 App 經實證可以減壓」。它能讀成的是——「在一群自我選擇的使用者裡,做完 session 後自評壓力會下降,這個訊號值得用對照組進一步驗證」。

最後,幫你整理一套資料素養的checklist。下次看到「樣本很大、效果量很漂亮」的健康或行為研究,照順序問五個問題:

第一,有對照組嗎?沒有,那大概只能講相關。第二,是前後測嗎?是的話,警惕迴歸均值跟期望效應。第三,怎麼測的?如果是自陳單題,那就很主觀、容易被期待影響。第四,樣本怎麼來的?自我選擇的話,不能外推到一般人。第五,這個大樣本,是在解決隨機誤差,還是被拿來掩蓋設計問題?

其實工程上是同一套直覺。A/B test 沒有 control,那就不叫 A/B test;觀測指標再漂亮,只要沒有對照、沒有隨機分流,你看到的常常是迴歸均值跟選擇效應,不是你那個改動的功勞。

所以這集留三個要點給你帶走:第一,大樣本不等於嚴謹,大效果量也不等於因果,這兩件事要分開看。第二,沒有對照組的前後測,最大的敵人是迴歸均值跟期望效應。第三,讀論文跟讀 dashboard,要小心的,其實是同一件事。

Tags

Related Articles