Psychedelics versus antidepressants: what does the 2026 meta-analysis tell us?
An in-depth analysis of the 2026 meta-analysis by Williams, Barnett and Szigeti, published in *JAMA Psychiatry*.
This is an expanded version of my Facebook post. If you arrived here from Facebook, you already know the punchline: 0.3 points on the HAM-D. Here, I go deeper—into the mathematics, the methodology and the questions this paper does not answer. If you have not read the Facebook post, no problem: it appears below as Part One. Part Two is the in-depth analysis I promised.
Already read the Facebook post? Jump straight to the in-depth analysis of the mathematics, methodology and unresolved questions.
Jump to the deep dive →Part One—the Facebook post
This meta-analysis has swept across the internet like a gale. Headlines are shouting, people are sharing and comment sections are multiplying. But when I look at these discussions, I mostly see reactions to the abstract. “Psychedelics no better than antidepressants!” Shock, disbelief, triumph or outrage. The holy war rolls on.
Yet an abstract is just that—an abstract. The real substance lies inside: in the methods, the discussion and what the authors leave implicit. That is where the meat of the paper is. And since fat carries flavour, this is going to be a substantial meal! :) Apologies to my vegan friends—the metaphor rather took on a life of its own ;-)
So, let’s dig in.
Remember how it began? Headlines such as “Psychedelics more effective than antidepressants”, “Psilocybin changes lives after a single dose” and “A revolution in psychiatry”. People shared them without a second thought. Who did not click “share” and think, At last? I will admit that I, too, was swept up in the excitement more than once.
To be fair, those headlines did not come from nowhere. Early studies of psilocybin, ayahuasca and LSD looked genuinely promising: large effects, rapid improvement and patients describing experiences that changed how they understood themselves. For someone who works with people experiencing depression, that is a powerful promise. For patients themselves, it may represent a breakthrough. The possibility that one profound psychedelic experience, held within a therapeutic relationship, could accomplish something that months of pharmacotherapy cannot is intellectually and clinically compelling. Is it not? It is hard not to become excited.
But excitement is not data. Still less is it evidence.
Recently, JAMA Psychiatry published a meta-analysis by Williams, Barnett and Szigeti (2026). A meta-analysis does not collect a new sample of participants. Instead, it synthesises the findings of multiple earlier studies. It is a statistical method that can reveal a pattern where any single study shows only a fragment. That is why well-conducted meta-analyses sit high in the hierarchy of evidence.
It is also why media outlets and influencers pounce on them—and why they deserve careful reading. A meta-analysis is only as strong as the studies it includes. Garbage in, elegantly packaged garbage out ;-) So, what happened this time?
This is a paper to which I expect to return. I read it about three times and still struggled to settle my thoughts about it. Whenever I found some peace, I lost it somewhere between admiration for its statistical elegance and discomfort with limitations that the authors sometimes treated as an afterthought. I think it is fair to call this analysis a masterpiece. But, like every masterpiece, it prompts questions that it cannot answer by itself.
It is not an easy read. There is plenty of statistics, methodology and mathematics: Bayesian modelling, ROPE, scale conversions and sensitivity analyses. You have to work through it. Once you do, however, you can see the beauty of the science. Yes, science is beautiful partly because of statistics. Perhaps especially because of statistics.
Enough introduction. The authors analysed 24 trials: eight of psychedelic-assisted therapy, including 249 patients in total, and sixteen of conventional antidepressants administered open-label, including 7,921 patients. It is worth pausing over those eight psychedelic trials, because the label “psychedelic-assisted therapy” covers very different interventions: psilocybin, ayahuasca, LSD and 5-MeO-DMT.
Each substance has a different pharmacological profile, duration of action and experiential character. The therapeutic protocols also differ: the number of preparatory sessions, the dose, the presence and role of the therapist, and the approach to integration afterwards. Pooling all of this and comparing it with antidepressants is a simplification that must be kept in mind when reading the results. A meta-analysis can reveal a pattern, but it also blurs differences.
The authors asked an elegantly simple question. What happens if we equalise the blinding conditions—or, more accurately, the unblinding conditions? What if patients in both groups know what treatment they are receiving?
That is the heart of the problem. In a conventional double-blind trial, neither the participant nor the researcher knows who has received the drug and who has received placebo. At least, allocation is intended to be difficult to infer. Escitalopram may upset your stomach, but it will not suddenly make the world more colourful.
This uncertainty about treatment allocation is the purpose of blinding. It helps reduce systematic differences in expectations and makes it easier to distinguish effects associated with the intervention from effects associated with believing that “this will help me”. In psychedelic trials, blinding often fails. After taking psilocybin, participants may not need to guess which group they are in. The experience is so distinctive that, in some trials, 90–95% correctly identify their allocation. In conventional antidepressant trials, the corresponding figure is around 60%. That said, some participants in psychedelic trials do not experience marked psychedelic effects even after a full dose, whereas some in a 1 mg group report such effects.
Blinding is not a technical footnote. It determines how confidently we can attribute an observed difference to pharmacology rather than to expectations or to the belief that one has received something groundbreaking. Functional unblinding can be mitigated, and some trials manage it better than others, but that is a subject for another post.
The authors of the meta-analysis took a clever approach. Rather than comparing psychedelics with placebo, where blinding is particularly difficult, they compared psychedelic-assisted therapy with antidepressants administered openly, without blinding. They attempted to “level the playing field” by comparing studies in which participants in both groups knew what they were taking.
The estimated difference on the 17-item Hamilton Depression Rating Scale (HAM-D or HAMD-17) was 0.3 points in favour of open-label antidepressants (95% CI, −1.39 to 1.98; p = .73). On a scale ranging from 0 to 52, the point estimate was clinically tiny, and the frequentist analysis found no evidence of a difference.
In my view, that is useful. Let the data cool our enthusiasm. This result does not show that psychedelics are ineffective. It shows no evidence that they outperform open-label conventional antidepressants under these conditions. Nor does a non-significant result, by itself, prove equality or non-inferiority; the Bayesian analysis discussed below addresses practical equivalence more directly. Enthusiasm unsupported by data can cause harm. It can lead patients to abandon a treatment that helps them in favour of one that currently promises more than it can demonstrate.
But the paper also has a less discussed side—one that only struck me on my third reading.
The average time to outcome assessment was 3.4 weeks in the psychedelic trials and 8.1 weeks in the antidepressant trials. That is a substantial difference: two frames taken from entirely different moments in the film. Conventional antidepressants had more time for their effects to develop. Psychedelic treatments were assessed earlier, before we knew what would happen to participants a month later. The authors acknowledge this limitation, but do not model time to endpoint statistically. We therefore do not know whether the two trajectories converge on the same destination or whether we are simply looking at different sections of the road.
Before anyone concludes, “So psychedelics do not work”, let us pause. That is not what this meta-analysis shows. It finds no evidence that psychedelic-assisted therapy is more effective than conventional antidepressants when one key factor—the participant’s knowledge of the treatment received—is made more comparable.
And here the more interesting discussion begins: what is actually therapeutic? The molecule, the emotional experience, the therapeutic relationship, hope—or all of these together, in proportions we do not yet understand?
Part Two—the deep dive
Fair warning: from this point onwards, the mathematics becomes denser. Bayesian modelling, ROPE, decomposition of effect sizes and questions about complex interventions. I promised this on Facebook. Here it is. I was on holiday, so I had time to read properly :)
Where did the “5-point advantage” come from?
To understand why this meta-analysis matters, we first need to understand what it is trying to explain. And to understand that, we must step back and consider how we determine whether an antidepressant works.
One instrument commonly used in clinical trials is the 17-item Hamilton Depression Rating Scale (HAM-D or HAMD-17). It is a clinician-rated scale, not simply a questionnaire completed by the clinician or patient. A trained rater assesses the severity of symptoms on the basis of an interview and clinical observation, applying defined scoring anchors to items concerning depressed mood, guilt, suicidal thoughts, insomnia, anxiety, somatic symptoms and psychomotor change. The ratings are summed to produce a total score ranging from 0 to 52; higher scores indicate greater symptom severity.
This brings us to a concept that will recur throughout the text: the minimal clinically important difference, or MCID. This is a threshold below which a difference may appear in the data but is unlikely to be meaningful in the patient’s life. For the HAM-D, a threshold of 3 points has been used by NICE (Moncrieff & Kirsch, 2015). Three points.
To get a feel for the scale, imagine two people with depression. One wakes at three in the morning and cannot return to sleep. The other sleeps poorly but makes it through the night. On the HAM-D, that difference might amount to two points on a single item. A three-point difference across the entire scale is around the point at which change may begin to matter clinically. These examples are illustrations, not literal translations of a score: the same three-point difference can arise from many different combinations of items.
Moncrieff and Kirsch also examined how large a HAM-D difference was needed before it corresponded to a clinician’s overall rating of improvement on the CGI-I. Their estimate was approximately 7 points. In other words, even the 3-point threshold lies below what a clinician is likely to register as a minimal global improvement.
It is worth pausing over something that is rarely said aloud. We have learnt to speak about depression through scales: so many points on the HAM-D, so many on the BDI, so many on the PHQ-9. This is useful because it gives us a shared language and allows studies to be compared. But depression is not a score. It is waking with the feeling that the day has no meaning before it has even begun. It is being unable to enjoy the presence of your own child. It is suffering that no rating scale captures in full. Perhaps one day we will develop methods that do so more successfully. For now, we have the measures we have, and we need to know how to read them.
How to read these numbers
After equalising unblinding, the psychedelic edge over placebo drops from 7.3 to 0.3 pts
below the threshold a patient can even feel
Where patients are · full HAM-D scale (0–52 pts)
Both groups, drug and placebo, improve by a dozen or more points. The whole debate is about the difference in that improvement, shown below.
How much treatment beats placebo · difference in HAM-D points
E.g. 3 points is roughly the difference between "can't get out of bed" and "can get up".
Back to the story. Historical trials of conventional antidepressants such as SSRIs and SNRIs found an average drug–placebo difference of approximately 2.4 HAM-D points. Two point four, against an MCID threshold of 3 points. This is part of the evidence base on which decades of depression pharmacotherapy have been built. It is not much, but it is what we have.
Then came the psychedelics. Trials of psilocybin, ayahuasca and LSD reported an average psychedelic–placebo difference of approximately 7.3 HAM-D points—roughly three times the 2.4-point difference reported for antidepressants. On paper, it looked like a qualitative leap. The headlines practically wrote themselves, even before ChatGPT existed.
The difference between those two estimates—2.4 and 7.3—is approximately 5 points. Zachary Williams, whose work includes metascience and clinical-trial methodology, Hannah Barnett and Balázs Szigeti, a neuroscientist known for his work on placebo effects and functional unblinding in psychedelic research, set out to explain its origin. Their paper, Williams, Barnett and Szigeti (2026), casts new light on earlier enthusiastic reports.
Mathematical deconstruction: where do those 5 points come from?
The authors decompose the apparent 5-point advantage into two components, in a way that is both fascinating and sobering.
Before describing them, an important caveat. The authors did not take a calculator and “prove” their thesis with a single equation. This is an interpretation based on comparing estimates from several analyses within their statistical framework. The arithmetic fits, but, like any interpretation, the explanation depends on assumptions and on the comparability of different sets of trials. Its strength lies in the coherence of the estimates, not in mathematical proof.
Component 1—the difference associated with open-label administration. In the conventional-antidepressant data, outcomes were approximately 1.29 HAM-D points better in open-label than in blinded trials. This estimate is consistent with an expectancy or blinding effect: knowing that one is taking an active medication may increase improvement. It is not, however, a pure experimental estimate of awareness alone, and the same open-label–blinded difference was not demonstrated for psychedelic-assisted therapy, which is functionally open-label even when a trial is nominally blinded.
Component 2—the nocebo effect, or “know-cebo”. Szigeti uses this neologism to distinguish classical nocebo—deterioration driven by negative expectations about a treatment—from a more specific possibility: deterioration or reduced improvement after a participant realises that they have not received the hoped-for active treatment. Placebo groups in psychedelic trials improved by approximately 4.0 HAM-D points less than placebo groups in conventional-antidepressant trials. The authors interpret this as consistent with a “know-cebo” effect. Because this is a comparison across different sets of trials, it cannot establish that placebo itself actively caused deterioration.
Imagine the situation. You have lived with severe depression for a long time and enrol in a trial. For weeks, you prepare for an experience that is supposed to change your life. You enter a dimly lit room, take a capsule, put on an eye mask and listen to music. An hour passes and nothing happens. You realise that you did not receive psilocybin. You received the comparator. The disappointment is unlikely to be neutral. It may affect both your expectations and the symptoms you report.
The two estimates—1.29 plus 4.0—sum to 5.29 HAM-D points, which closely matches the 4.9–5.0-point difference between the antidepressant–placebo and psychedelic–placebo contrasts in the traditional comparisons. That numerical fit is striking. It supports the coherence of the authors’ explanation, although it does not prove that these two mechanisms causally account for every part of the difference.
The 4.0-point difference between the placebo groups alone amounts to approximately 55% of the 7.3-point psychedelic–placebo contrast. In other words, much of the apparent advantage may reflect how the comparator groups behave under functional unblinding rather than superior pharmacological efficacy. In the authors’ interpretation, it is not simply psychedelics winning; it is also placebo losing. They argue that the apparent superiority of psychedelics in blinded trials is substantially influenced by functional unblinding.
Deconstructing the 5-point advantage
The observed advantage of ~5 points fits entirely within methodological artifacts. Expectancy + know-cebo = 5.29 pts. How much of this is actual pharmacology?
¹ In a clinical trial, “placebo” is not always an inert sugar pill. Psychedelic studies have used inert placebos such as cellulose or lactose, active comparators intended to make allocation less obvious, and very low doses of the study drug itself, such as 1 mg of psilocybin. Each control condition affects blinding and participants’ experiences differently.
² Nicotinamide, also known as niacinamide, is the amide form of vitamin B3. It was used at a dose of 100 mg as the active comparator in the EPISODE trial. It should not be confused with nicotinic acid, or niacin: unlike nicotinic acid, nicotinamide does not normally cause flushing, itching, burning or tingling. Whatever allocation cues it may provide are qualitatively unlike a psychedelic experience.
Bayesian analysis: how certain can we be?
The authors did not stop at frequentist statistics. To understand why that matters, it is worth pausing over how the two approaches differ.
A frequentist analysis asks: “If there were no true average difference between psychedelic-assisted therapy and conventional antidepressants, how unusual would a result at least as extreme as the one observed be?” That is what the p value addresses. Here, p = .73, so the observed estimate is entirely compatible with the null hypothesis of no mean difference. We therefore have no basis for rejecting that null hypothesis. This does not, by itself, establish that the treatments are equal. Failure to detect a difference is not proof of equivalence.
A Bayesian analysis asks a different question. Given the data, the model and the prior assumptions, how probable are different values of the true difference? This distinction is subtle but fundamental. Frequentist inference describes the probability of data, or more extreme data, under an assumed hypothesis; Bayesian inference describes uncertainty about parameters or hypotheses conditional on the data and model. The latter is often closer to the question we intuitively want answered.
The authors defined a ±3-point HAM-D interval as the region of practical equivalence (ROPE), based on the 3-point MCID. A true difference inside that interval would be regarded as too small to be clinically important on average.
Their Bayesian analysis estimated a posterior probability of just 0.2% that psychedelic-assisted therapy was superior to conventional antidepressants by at least 3 HAM-D points. Put loosely, if we made 500 bets on clinically meaningful superiority under this model, we would expect to win about one.
At the same time, 99.1% of the posterior distribution lay within the ROPE of ±3 HAM-D points. Think of the ROPE as a band around no effect. If almost the entire posterior distribution lies inside it, the data and model support the conclusion that any average difference is unlikely to exceed the prespecified threshold of clinical importance.
This is stronger than simply saying, “We did not detect a difference.” It is positive Bayesian evidence of practical equivalence within the specified ±3-point margin, under the model and prior assumptions. It is not proof that the true difference is exactly zero. Celebrate or cry? Perhaps both, depending on what we expected.
The time problem: 3.4 versus 8.1 weeks
I flagged this issue in Part One, but want to examine it more closely because I consider it the most serious limitation of the meta-analysis.
First, some context. In a clinical trial, an endpoint is the outcome assessed at a prespecified time. The primary endpoint is the outcome on which the study’s main conclusion is based.
The mean time to assessment of the primary endpoint was 3.4 weeks in the psychedelic trials and 8.1 weeks in the antidepressant trials. SSRIs often require several weeks for their full clinical effect to emerge, so assessment at eight weeks is more likely to capture an established response. Psychedelic-assisted therapy was assessed at 3.4 weeks, when changes related to the experience and its integration may still have been unfolding.
Comparing these endpoints is like comparing a sprinter with a marathon runner at the one-kilometre mark. The sprinter looks impressive, but we do not yet know who will reach the finish line.
The authors acknowledge the discrepancy in their limitations section and note that it may have affected the results. They do not, however, include time to endpoint in the statistical model. It remains a caveat rather than an analysed source of variation.
This matters because methods exist for modelling trajectories of change rather than only static endpoints measured at different times.
Something personal stopped me here. In 2018, I published a paper with Ludmiła Zając-Lamparska and Monika Deja on latent growth curve modelling (LGCM) in Polskie Forum Psychologiczne. At the time, I was not thinking about psychedelic research at all; that entered my professional life two years later. Yet the problem of time in longitudinal data turns out to be highly relevant to psychedelic psychiatry.
Latent growth curve models analyse longitudinal data as trajectories rather than as isolated observations. They estimate the starting point, or intercept, and the rate of change, or slope, separately and, most importantly, model individual differences in those trajectories. Instead of asking only “How much did it change?”, we can ask “How did the change unfold, and for whom did it unfold differently?”
What if, rather than comparing endpoints assessed at different times without accounting for timing, the authors had examined time to endpoint as a study-level moderator—for example, by using meta-regression? This would not reconstruct individual trajectories, but it could test whether timing was associated with the estimated effect. The picture might have looked different. I am not saying that the result would necessarily have changed. I am saying that it would have been more complete.
Individual participant data (IPD) meta-analyses can model change curves at the level of individual participants rather than relying on aggregated study results. They require access to raw data from every study, which is logistically difficult but methodologically possible. We do not know whether the trajectories of people treated with psychedelics and antidepressants converge on the same destination, or whether we are simply looking at different parts of the road. This is where the field should go next.
Sprinter vs marathoner
Comparing response times of psychedelics and antidepressants in clinical trials
We don't know what happens after 3.4 weeks. Every line on the chart is a possible trajectory. We know none of them. Comparing a sprinter with a marathon runner at the 1 km mark doesn't tell us who will reach the finish line. We need longer studies.
Note: this visualization reflects measurement windows from the Williams et al. (2026) meta-analysis. Individual trials use different time points.
The EPISODE trial: what does it tell us about mechanism?
The Williams, Barnett and Szigeti meta-analysis found no evidence that psychedelic-assisted therapy was more effective than open-label antidepressants under more comparable unblinding conditions. But it does not tell us which components of psychedelic-assisted therapy are responsible for change. That is why it is worth considering another study published in the same issue of JAMA Psychiatry.
The EPISODE trial (Mertens et al., 2026) was a two-centre, triple-blind, phase 2b randomised clinical trial involving 144 people with treatment-resistant depression. It compared 25 mg psilocybin with 5 mg psilocybin and an active comparator, 100 mg nicotinamide. Participants received two administrations six weeks apart according to one of four dosing sequences.
The primary endpoint was treatment response—at least a 50% reduction in HAMD-17 score—six weeks after the first administration. Response rates were 17.0% with 25 mg psilocybin, 12.5% with 5 mg and 10.6% with nicotinamide. The prespecified comparison between 25 mg psilocybin and nicotinamide was not statistically significant (adjusted OR 1.73, 95% CI 0.53 to 6.23; p = .19). This was an inconclusive primary result, not evidence that the interventions were identical. Key secondary outcomes nevertheless provided exploratory evidence of clinically meaningful reductions in depressive symptoms with 25 mg psilocybin.
The study contained another intriguing observation. Higher scores on the Emotional Breakthrough Inventory—reflecting experiences of emotional openness, insight or an inner turning point—were associated with better depression outcomes. The authors interpreted this association as supporting a possible mechanistic role for the quality of the acute psychedelic experience. It remains an association: it does not establish that emotional breakthrough caused the improvement or that its contribution can be separated from dose, expectancy and functional unblinding.
This gives me pause. If the quality of the emotional experience contributes to antidepressant change, then we are speaking about something more complex than “a drug for depression”. We are speaking about a complex intervention.
What is a complex intervention? The substance, the therapy, the setting or the hope?
Psychedelic-assisted therapy is not simply a pill. It combines a substance with therapeutic preparation, the presence of a therapist—often two—for six to eight hours during the experience, and integration sessions afterwards. As Muthukumaraswamy and colleagues (2025) argue, it is a complex intervention in which the substance, psychotherapy, setting and the patient’s expectations are intertwined.
The randomised clinical-trial paradigm was designed to test the isolated effect of an intervention: separate the drug from its context, give some participants the drug and others placebo, and estimate what the drug itself does. In psychedelic-assisted therapy, that separation may be structurally impossible. This is not a matter of researcher incompetence; the object of study may itself be indivisible.
There is a tool called PRECIS-2—the Pragmatic–Explanatory Continuum Indicator Summary—that assesses where a clinical trial lies between the explanatory pole, which tests efficacy under controlled conditions, and the pragmatic pole, which tests effectiveness in routine practice. Psychedelic trials have been rated at approximately 15–16, placing them close to the explanatory pole and far from the pragmatic one. We study psychedelics in conditions involving a dimly lit room, music, two therapists and a session lasting many hours—conditions that may bear little resemblance to routine psychiatric care. This is interesting, but it is also a problem. PRECIS-2 deserves its own post.
What do the experts say?
James Rucker, a psychiatrist and senior lecturer at the Institute of Psychiatry, Psychology & Neuroscience at King’s College London, leads the Psychoactive Trials Group and is one of Europe’s most experienced clinical researchers in this field. Commenting on the two papers for the Science Media Centre, he made a point worth remembering: the findings are consistent with an antidepressant effect of psychedelic-assisted therapy, but suggest that positive expectations may account for a meaningful part of the observed change. He added an important clinical point. In routine practice, medication is not administered blindly; patients know what they are taking. An open-label comparison may therefore resemble the consulting room more closely.
David Owens, emeritus professor of clinical psychiatry at the University of Edinburgh, emphasised that both papers touch on the central problem in evaluating psychedelics as therapeutic tools: blinding. Until that problem is addressed, interpretation will continue to go round in circles.
Commentators discussing EPISODE also noted the divergence between its non-significant primary outcome and exploratory evidence from secondary measures. The published report found no additional benefit from a second 25 mg administration within the study period; by week 12, after every participant had received 25 mg psilocybin at least once, no between-group differences remained, although the second treatment phase was not powered to detect such differences. Whether longer-term outcomes depend mainly on pharmacology, the acute experience, continuing therapeutic engagement or some combination of these remains an open question.
What does this meta-analysis leave unresolved?
The Williams, Barnett and Szigeti meta-analysis is elegant, but it is not omnipotent. These are the questions that continue to trouble me and that the paper cannot answer.
Does practical equivalence on the HAM-D imply broader clinical equivalence? The Hamilton scale measures symptoms such as insomnia, appetite loss and psychomotor retardation. It does not directly measure meaning, relationships, occupational functioning or quality of life. In the comparison of psilocybin with escitalopram (Carhart-Harris et al., 2021; Erritzoe et al., 2024), the primary depression outcome did not differ significantly, while some secondary measures of functioning and well-being favoured psilocybin. Equal or similar symptom scores can therefore coexist with different clinical realities. I take this argument further in Invisible variables, where I ask what a scale captures and what remains outside its frame.
Is episodic administration clinically preferable to continuous treatment? If one psychedelic experience produces a similar average reduction in HAM-D scores to months of daily SSRI treatment, without continuous exposure to side effects such as emotional blunting, weight gain or sexual dysfunction, that similarity may look very different from the patient’s perspective. Practical equivalence on one symptom scale is not equivalence of the treatment experience.
What about the trajectory over time? The 3.4-versus-8.1-week problem is not merely technical. It concerns whether psychedelic-assisted therapy reaches its maximum effect more rapidly and whether that effect persists. The meta-analysis does not answer either question.
What about treatment-resistant depression? The authors performed a sensitivity analysis excluding trials involving treatment-resistant depression, and the result did not materially change. But that does not answer whether psychedelics may be a valuable option for an individual who has not responded to a third conventional antidepressant, even if average efficacy is similar at population level. For that person, “not better on average” may still mean “a remaining option”.
Where does this leave us?
This meta-analysis is important and necessary. It does precisely what it should: it cools a narrative that had raced ahead of the data. But rather than closing the discussion, it opens it again.
Psychedelic-assisted therapy probably has antidepressant effects. The question is what produces them: the molecule, the emotional experience, the therapeutic relationship, hope, or all of these together in proportions we do not yet understand.
I will be speaking about this at the Nauka Psychodeliczna 2026 conference on 13–14 June at the University of Warsaw, in a talk entitled “Invisible Variables: What We Still Don’t Know About Psychedelic Therapy”. I will focus on clinical application—what these data mean for the person sitting across from me in the consulting room, not only for a statistical model.
If this subject interests you, register here: naukapsychodeliczna.org/etn/konferencja-nauka-psychodeliczna-2026
Curious? :)
References
Carhart-Harris, R. L., Giribaldi, B., Watts, R., Baker-Jones, M., Murphy-Beiner, A., Murphy, R., Martell, J., Blemings, A., Erritzoe, D., & Nutt, D. J. (2021). Trial of psilocybin versus escitalopram for depression. New England Journal of Medicine, 384(15), 1402–1411. https://doi.org/10.1056/NEJMoa2032994
Erritzoe, D., Barba, T., Greenway, K. T., Murphy, R., Martell, J., Giribaldi, B., Timmermann, C., Murphy-Beiner, A., Baker-Jones, M., Nutt, D. J., Weiss, B., & Carhart-Harris, R. L. (2024). Effect of psilocybin versus escitalopram on depression symptom severity in patients with moderate-to-severe major depressive disorder: Observational 6-month follow-up of a phase 2, double-blind, randomised, controlled trial. eClinicalMedicine, 76, 102799. https://doi.org/10.1016/j.eclinm.2024.102799
Mertens, L. J., Koslowski, M., Betzler, F., et al. (2026). Efficacy and safety of psilocybin in treatment-resistant major depression: The EPISODE randomized clinical trial. JAMA Psychiatry, 83(5), 448–460. https://doi.org/10.1001/jamapsychiatry.2026.0132
Moncrieff, J., & Kirsch, I. (2015). Empirically derived criteria cast doubt on the clinical significance of antidepressant-placebo differences. Contemporary Clinical Trials, 43, 60–62. https://doi.org/10.1016/j.cct.2015.05.005
Muthukumaraswamy, S. D., Baggott, M. J., Schenberg, E. E., Decker, R., & Reckweg, J. T. (2025). Psychedelic-assisted therapy as a complex intervention: Implications for clinical trial design. Therapeutic Advances in Psychopharmacology, 15, 20451253251381074. https://doi.org/10.1177/20451253251381074
Williams, Z. J., Barnett, H., & Szigeti, B. (2026). Psychedelic therapy vs antidepressants for the treatment of depression under equal unblinding conditions: A systematic review and meta-analysis. JAMA Psychiatry, 83(5), 461–468. https://doi.org/10.1001/jamapsychiatry.2025.4809
Zając-Lamparska, L., Warchoł, Ł., & Deja, M. (2018). Analiza danych podłużnych: Modelowanie latentnych krzywych rozwojowych [Longitudinal data analysis: Latent growth curve modeling]. Polskie Forum Psychologiczne, 23(2), 395–412. https://doi.org/10.14656/PFP20180210