The best evidence Tylenol causes autism isn't great
On Monday, RFK Jr announced Tylenol ‘causes’ autism referencing three studies as evidence. But if this is there best evidence against Tylenol, it leaves a lot to be desired.
Let’s talk about the biggest study: a review paper published about a month ago. If you’ve been reading E is for Epi for a while, you know the first step in assessing a paper is to get the study. So here’s the link: https://ehjournal.biomedcentral.com/articles/10.1186/s12940-025-01208-0
Next, let’s take a minute to look at the authors and where this was published. Remember, we’re looking for red flags here, not trying to make an argument from authority. The research team for this paper is four environmental health researchers, from reputable institutions, including the Dean of the Harvard TH Chan School of Public Health. I’m not going to dig into their backgrounds,1 but I do want to note that it’s not clear any of them are either pharmacoepidemiology or perinatal epidemiology experts and those are the two specialities which are relevant to this research question. It’s also worth noting that this review paper is published in a journal called Environmental Health. Again, a good journal, but not focused on either pharmacoepidemiology or perinatal epidemiology.
This isn’t necessarily a red flag, but it’s maybe an orange one. At minimum it raises two important questions: (1) Why didn’t they collaborate with any experts in the specific area of drug exposures during pregnancy? (2) Why didn’t they publish in a journal focused on (and therefore reviewed by experts in) that specific area?
So, with that out of the way, the next step is to look at the title and abstract.
Title: “Evaluation of the evidence on acetaminophen use and neurodevelopmental disorders using the Navigation Guide methodology”
The first thing to note from the title: this study is not restricted to autism as an outcome, rather they are looking at all neurodevelopmental disorders.
The second thing to note: they are using the “Navigation Guide methodology”
These are the two key pieces of information the authors wanted to include in the title, so they set the tone for our understanding of this study. The first tells us that we are going to need to read carefully if we only want to know about acetaminophen and autism. Because there will be information about other neurodevelopmental disorders that we aren’t necessarily interested in right now.
The second tells us what we should expect from the methods: that they follow something called the “Navigation Guide”. But what is the Navigation Guide? It turns out it’s an “Evidence-to-Decision Framework for Environmental Health”. Which explains why I’d never heard of it before, and why most of the epidemiologists I’ve talked to in the past 4 days have never heard of it before either. It’s not a general framework for evaluating evidence. It’s specific to environmental health. Which is the expertise of the authors. But is not the topic of the review paper.
There are plenty of other frameworks for guiding systematic literature reviews. Even some that are specifically tailored to evaluating questions about whether or not causal effects exist when the only available evidence is from observational data. Ideally, the authors should have given some justification as to why the chose the Navigation Guide and not one of the other options available to them. Unfortunately they did not.
Abstract: The abstract tells us that this is a review of the scientific literature on prenatal acetaminophen exposure and neurodevelopmental disorder.
They give us the total number of studies they found (46) and tell us numerically how many studies showed positive, negative, or no relationship between acetaminophen and any neurodevelopmental disorder.
The abstract does not tell us how many studies they found on acetaminophen and autism, nor what those specific findings were.
The abstract conclusions are very strong: they advise that pregnant women immediately limit acetaminophen use.
I actually read this entire study in the order it was written, because I wanted to make sure that I didn’t miss anything. But the next step for me usually is to go straight to the methods section, so that’s what I’m going to do here. I have a lot that I could say about the methods section of this paper, but I’ll try to keep it short:
Based on the information the authors provide, their search process is insufficiently broad and is likely to miss studies that found no association.
This is always a concern in systematic literature reviews and is why the search and identification stage is always so time consuming. Because papers which find no association are much less likely to be published. And so you have to dig much deeper to find them. These authors didn’t.
The authors quantify study quality in 3 ways: “risk of bias”, “strength of evidence”, and “expert opinion”. They only explain “risk of bias” and don’t tell us anything about how they determined the other 2 scores.
This is a major lack of detail on the key information they ultimately use to synthesize the evidence they find and to draw their conclusions.
Interestingly, the peer-review reports of this study are public, so we can see that one of the peer reviewers did indeed ask for more detail on these scores.
The authors responded by adding text to the results discussing the score values but no new methods text. And we know this because the response to reviewers is also public (see Reviewer 2, comment 24).
Lastly, it’s worth mentioning that the methods section includes several contradictory statements (for example, saying in one place that meta analyses were included but in another that they were excluded), as well as references to a figure which doesn’t exist (or rather, which was clearly revised without updating the text).
Figure 1 in the manuscript is the overall study selection figure. But in the text they reference it as if it had three separate sections, one for each outcome of interest (ADHD, autism, and ‘other neurodevelopmental disorders’).
The methods section bottom line is this: There is insufficient information provided to allow anyone to reproduce what they have done with any confidence.
Let’s turn next to the results related to the autism outcome specifically (Table 2). There were 46 studies identified overall but only 8 of them are listed as telling us anything about autism. And, if you look closer at Table 2, you’ll see that in fact there were only 7 studies; one study is listed twice in the table (Ahlqvist et al 2024), with separate rows for the “overall study” and for the “sibling-controlled analysis”.2
The next thing worth looking at in Table 2 is the “Outcome definition” column. Here, you’ll see that the outcome ranges from “mother’s self-report” to “autism from ICD codes”. That’s a wide range, but it’s not unreasonable.
But you might also notice that the outcome for the 3rd row (Alemany et al, 2021) is listed as “subscale of the CBCL11/2-5 and CBCL6/18 and ADHD criteria of DSM-IV)”, which at first read sounds like it’s actually a measure of ADHD. The CBCL is the “child behavior checklist” and the numbers indicate the age-range version of this checklist (1.5-5 years; 6-18 years). There are multiple different subscales of this checklist so it’s a bit strange that they don’t mention which one. But, regardless of which subscale they used, this seems to be essentially a “related symptoms” outcome and not a “diagnosis” outcome. That’s not ideal.
One final thing you might notice looking through Table 2 is that the “Risk Estimate” provided for the 4th row (Aveila-Garcia et al 2016) says explicitly that it is for ADHD. Is that a typo? Or did this study not look at autism? We’ll have to go read it ourselves to find out. And we will. But not today.
For now, let’s move on to their evaluation of these 7-ish studies.
Table 6 gives us their assessment of the “risk of bias” of these studies by category of type of bias considered. The values range from 1 (low risk) to 4 (high risk) for each individual study and category. Most of the table is 1s, indicating that they felt that for most issues, most studies were low risk of bias. Let’s look at the 4s:
The Ahlqvist study is categorized as “high risk of bias” for exposure assessment, both for the overall study and for the sibling-controlled analysis.
One other study gets a 4 here, and most of the rest get 2 (“probably low risk of bias”). So how do they assess exposure?
Ahlqvist et al get exposure information by interviewing midwives during the pregnancy.
This is Sweden, and I don’t know much about how common midwives are there but presumably they are common enough for this to have been a feasible data gathering technique for 2.4 million births!
Without further explanation, there doesn’t seem to be any reason to assume these midwives would give incorrect information. 4 seems inappropriate.
The other “4” for exposure information is a study which used maternal self-report after autism diagnosis. That approach does have a high risk of bias, specifically a type of bias called recall bias. A 4 is appropriate here.
The studies coded as 2 (probably low risk) all obtained exposure information by interviewing the mothers during pregnancy.
Why is interviewing mothers during pregnancy “probably low risk” but interviewing midwives during pregnancy “high risk” for bias? These don’t seem that different to me.
The only “low risk” exposure assessment is a study which measured acetaminophen in cord blood. But that only gives a snapshot of exposure at the end of pregnancy, and isn’t necessarily as useful as exposure measured during pregnancy. It’s not clear to me that this is necessarily low risk of bias.
The rest of the table is fairly uniform, except that the sibling-controlled analysis in the Ahlqvist study is categorized as “probably high risk of bias” for confounding, which is a strange choice since the use of sibling controls is specifically because this design limits the impact of confounding. In the text, they say that this is because sibling-controlled studies could be affected by sibling-level confounding. And that’s true, but there were over 16k siblings in this study which is a large enough sample to control for within-sibling confounding if it’s a problem. Did they? This paper doesn’t tell us.
Finally, there’s also a column for “other problems that could put it at a risk of bias” where the Ahlqvist study again is scored very poorly. Again, it’s not clear why.
Last but not least, is Table 7. This table scores each study on strength of evidence, with columns (a) through (f) presumably contributing to the overall “strength of evidence” score (although the overall score clearly isn’t the sum or the average). It also gives us an “expert opinion score” but we aren’t told where this comes from nor who the experts are.
The Ahlqvist study again scores the worst in this table, and if we look at the individual categories we see that is in large part because of category (b) “Large effect (>2)”. At this point it’s worth noting that the Ahlqvist study finds no effect of Tylenol on autism (the overall study finds a very small effect, and this disappears in the sibling-controlled analysis).
That’s right: They down-weight the quality assessment of the largest study because it’s results show no relationship between Tylenol and autism.
Sit with that a minute.
They say this study is “weak/no evidence of an effect” because it’s results are null. And as a result, they discount it in their synthesis.
But based on everything they’ve told us about this study3, it sounds a lot more like it’s actually strong evidence of no effect. This review doesn’t even seem to consider this a possibility.
And that’s really the overall bottom line here.
This study doesn’t evaluate whether or not acetaminophen causes autism. It assumes that it does. And then it looks to see how strong the evidence supporting that claim is.
But in the end, that evidence is really not very strong.
Excluding Ahlqvist (which, after all says there is no association), there are at most 6 studies on acetaminophen and autism. If we just rely on the information they tell us about these studies in Tables 2, 6 and 7, at least one of them might not even have looked at autism at all and another study maybe only have looked at symptoms, not diagnoses. That leaves us with maybe 4 studies. And the authors themselves score 2 of these studies as such low quality that they don’t even included them in their final assessment table (Table 7).
So, we’re essentially left with only 2 studies that both assess the question of interest and meet the authors’ own quality criteria. That’s few enough studies that we can read them ourselves, and not rely on the opinion of this paper’s authors.4 And it’s hardly a smoking gun.
Tl;dr: if this paper is the best evidence that the US government could find that acetaminophen causes autism, then it’s fair to say there is not very good evidence at all.
Okay, that’s all for now. Next time, we’ll look at the other two studies RFK Jr mentioned in the press conference. Subscribe so you don’t miss out!
As an independent epidemiologist, I make my living entirely through reader support. Readers like you make E is for Epi possible. If you find value in this work, consider becoming a paid subscriber.
If you’re curious, there are quite a few articles out there discussing the publication history of the authors and conflicts of interest regarding paid expert witness testimony.
Incidentally, this also violates their methods since they say several times in the methods section that they made sure to select only the largest and/or most rigorous study when there were more than one available for a given dataset. They do not justify including the Ahlqvist study twice.
Don’t worry, I’m reading the Ahlqvist study too and in a future article will give you my review of it!
And we definitely will. But not today.




Critiques like this are valuable — not because they settle the debate, but because they remind us what careful science looks like: working through methods, assumptions, and red flags.
But here’s the thing: we should be doing this with every study. Whether it comes from Harvard, the NIH, or a preprint server, the same scrutiny should apply. Too often, “trust the experts” turns into a free pass for some research, while other equally credentialed (i.e. from a prominent researcher or institution) work gets torn apart the moment it crosses a political fault line.
That behind-the-scenes process — the messy, unglamorous scrutiny that is science — is what I find most worth writing about.
I think it's more than a little suspicious that this paper was published in August of this year, in a smaller, off-topic journal (that I'm guessing has a quicker turnaround), without any input from subject matter experts in the field. Almost like someone requested it? And a lead author is a Dean at Harvard at the same time that Harvard was appealing a huge freeze on federal funding to their institution. What do we know about the leadership of this journal? Any conflicts there?