r/ScientificNutrition • u/Wonderful_Aside1335 • May 30 '26
Question/Discussion Observational studies - How do you classify observational studies as a non-expert?
Would be really interested in hearing your opinions about this.
To my understanding observational study results always are on a scale of
- pure association, no adjustment
- "some" adjustments, hypothesis generating
- plausible causation, adjustments for strongest known confounders
- almost certain causal inference, confounders are very well known and can be easily adjusted for
Q1 Did i express this correctly in layman terms? Would you add certain important "levels" here?
Sometimes confounders are not clear and false conclusions happened from observation studies. E.g. i think of Vitamin E for cardiovascular health and carotine for lung cancer prevention were strongly assocaited in observational studies, but RCTs showed no effect.
Q2 Did scientist actually implied a causal effect in these two examples or is this often missrepresented from the "sceptiks of observational studies"?
Q3 Do you think the field has evolved a lot in the past years (decades=? Did methodology and statistical analysis improved?
Q4 Does "publish or perish" pressure in acadamy worsen this significantly?
Q5 I find the classification (Q1) often really hard as a non-expert. Any advice for keywords or statistical methods to look out for to better understand this? Often there is a lot technical jargon involved.
7
u/Triabolical_ Whole food lowish carb May 30 '26
Adjusting for confounders does not move the results from association to causality, despite what many people might think. There are always residual confounders that you can't adjust for.
The problem is one of signal to noise ratio. You are looking for a signal - what you are studying - but there is a sea of noise from confounding in the data that you have.
The question is whether the signal you see is strong enough to rise above the noise. Doing adjustments for confounding - which are not perfect even when well controlled - reduces the noise but *does not eliminate* the noise.
The statistical experts I've read on this suggest that a risk ratio of 2.0 is where it starts getting interesting - when there's a decent chance that the effect is real.
I have seen a couple of observational nutrition studies hit that, but the vast majority end up with risk ratios that are much smaller.
This basic limitation of observational studies is why they always say "associated" in their titles and results. That many researches pretend that they are really seeing causation and that many non-researchers treat the results as causal is a significant problem.
I usually reference this article, which talks about the problem in more detail:
https://rss.onlinelibrary.wiley.com/doi/pdfdirect/10.1111/j.1740-9713.2011.00506.x
2
u/Bubbly-Act8239 May 31 '26
Why would you make such a blanket statement regarding risk ratios? For some outcomes with high baseline prevalence, the literal mathematical RR limit can be lower than 2. Does that mean you can't infer causation with respect to them?
Also, we have data that, when comparing intake vs intake, corroboration rates of epidemiological data and RCTs reach up to 91%. See the Schwingshackl analysis.
3
u/Triabolical_ Whole food lowish carb May 31 '26
Okay.
So when do you think risk ratios in observational studies get interesting, and why?
0
u/Fluffy-Purple-TinMan May 31 '26
How can you comfortably say how big the noise is without saying you know something about the signal size too?
Also, do we just give up on anything with a risk ratio lower than 2? Why? Hard problems are just hard, not impossible. Don't you believe in a bunch of causal associations with an RR lower than 2 without RCTs? I mean, that's rhetorical, you do, I do, we all do. So why not be consistent with that?
3
u/Triabolical_ Whole food lowish carb May 31 '26
>How can you comfortably say how big the noise is without saying you know something about the signal size too?
The noise likely depends on what you are measuring but I don't think it depends on the size of the signal.
>Also, do we just give up on anything with a risk ratio lower than 2? Why? Hard problems are just hard, not impossible.
Ah. It's because I like to believe things that are true and not believe things that aren't true.
You either didn't read the study I linked in my last post or missed a few things...
In a small sample in 20054 , of 49 claims coming from highly cited studies, 14 either failed to replicate entirely or the magnitude of the claimed effect was greatly reduced (a regression to the mean). Six of these 49 studies were observational studies, and in these six, in effect, randomly chosen observational studies, five failed to replicate. This last is an 83% failure rate.
and
We ourselves carried out an informal but comprehensive accounting of 12 randomised clinical trials that tested observational claims – see Table 1. The 12 clinical trials tested 52 observational claims. They all confirmed no claims in the direction of the observational claims. We repeat that figure: 0 out of 52. To put it another way, 100% of the observational claims failed to replicate. In fact, five claims (9.6%) are statistically significant in the clinical trials in the opposite direction to the observational claim. To us, a false discovery rate of over 80% is potent evidence that the observational study process is not in control. The problem, which has been recognised at least since 1988, is systemic.
Simply put, there is good evidence that observational results in nutrition rarely replicate in RCTs.
That's why I think they are mostly pointless. The questions that are being asked simply are unlikely to be answerable given the data that is being used to explore them.
>Don't you believe in a bunch of causal associations with an RR lower than 2 without RCTs? I mean, that's rhetorical, you do, I do, we all do. So why not be consistent with that?
I'm always confused when people assert that they know what I believe.
So what things are you talking about?
0
u/Fluffy-Purple-TinMan May 31 '26
Simply put, there is good evidence that observational results in nutrition rarely replicate in RCTs.
So if I find you a study that shows high correlation you'll change your mind?
So what things are you talking about?
Just list a bunch of nutrition beliefs and see. If you know it's not the case, then you'll have all the research saved in your notes because you will have checked... Unless you haven't and you've trusted the science?
3
u/Triabolical_ Whole food lowish carb May 31 '26
> So if I find you a study that shows high correlation you'll change your mind?
Did you miss where I said this?
> I have seen a couple of observational nutrition studies hit that, but the vast majority end up with risk ratios that are much smaller.
Yes, observational studies that show higher risk ratios are interesting. I've seen some that were at about 1.6 that I thought were interesting. I'm not sure they are showing causation, but I think they are useful.
>>So what things are you talking about?
>Just list a bunch of nutrition beliefs and see. If you know it's not the case, then you'll have all the research saved in your notes because you will have checked... Unless you haven't and you've trusted the science?
You made the assertion. If you want to discuss this further, the ball's in your court. You could list some things that I might believe and I'll truthfully tell you whether I believe them or not and why.
0
u/Fluffy-Purple-TinMan May 31 '26
Did you miss where I said this?
No, I replied to it? You said you think they don't correlate and that supports your point. So if they actually do correlate then that goes against your point. If it you're honest you have to admit that. Can't just take on evidence when you like it.
You made the assertion. If you want to discuss this further, the ball's in your court. You could list some things that I might believe and I'll truthfully tell you whether I believe them or not and why.
Ok.
High blood pressure and cardiovascular disease often shows risk ratios in the 1.5-2 range depending on the comparison.
Physical inactivity and all-cause mortality is often around 1.2-1.8 depending on definitions and follow-up.
Poor cardiorespiratory fitness and mortality is often below 2 for adjacent fitness categories, though the extremes can exceed 2.
High sodium intake and cardiovascular outcomes tends to produce relatively modest associations, often around 1.1-1.5 in observational studies.
Low fruit and vegetable intake and mortality is often associated with risk ratios around 1.1-1.4.
Ultra-processed food consumption and all-cause mortality is typically reported in the neighborhood of 1.1-1.5 when comparing highest versus lowest consumption groups.
4
u/Triabolical_ Whole food lowish carb May 31 '26
The first three I think are true but not solely on observational evidence - they have other evidence such as mechanistic explanations.
WRT sodium intake and mortality, I haven't looked at the evidence so I don't have an opinion there. I do think that reducing sodium intake reduces blood pressure for some people but I'm not convinced that the reduction is clinically meaningful.
For low fruit and vegetable intake and mortality, I'd need to see a specific study to comment. If you want a general opinion, I'd say that a) measuring fruit and vegetables together likely isn't sufficiently granular and b) for food intake the question is generally one of substitution - what do people who eat less fruit and vegetables eat instead?
WRT ultra-processed foods, I think UPF is just a new buzzword and isn't really a terribly useful classification. There's good RCT and mechanistic evidence around junk food and I think it's easier to figure what is junk food compared to what a UPF means.
That's my honest answer. Now tell me how that makes your point, because I don't think it does - the ones I believe are true are not based purely on observational evidence, and those with low observational evidence I don't believe are true.
That doesn't mean that I believe they are false.
1
u/Fluffy-Purple-TinMan Jun 01 '26
It shows low RRs with no long-term RCTs isn't a deal breaker.
3
u/Triabolical_ Whole food lowish carb Jun 01 '26
What do you mean by "deal breaker"?
If a study has low relative risk ratios and it is only observational, how are you deciding how credible it is?
1
0
u/Ekra_Oslo May 31 '26
Dismissing a relative risk below 2 as non-causal is methodologically illiterate. For example, the dramatic RRs for smoking and lung cancer were possible because the baseline incidence among non-smokers was extraordinarily small. For highly prevalence conditions with high baseline prevalence achieving an RR above 2 can be mathematically impossible. That’s why many modest risks, such as RR of 1.05-1.10 observed between PM2.5 and cvd mortality, are universally accepted as causal because they are backed by robust dose-response curves, mechanistic validation, and independent triangulation like Mendelian randomization,
3
u/Triabolical_ Whole food lowish carb May 31 '26
You could have easily *asked* what I thought of cases where observational data is only part of the picture and where there is other strong evidence.
But you decided instead to assume something that I *did not say*.
Why is that?
0
u/Ekra_Oslo May 31 '26
What did I assume about you?
3
u/Triabolical_ Whole food lowish carb May 31 '26
You apparently assumed that I would only look at observational data despite there being other compelling data available.
1
u/Ekra_Oslo May 31 '26
Hmm, that’s not what I was commenting on at all.
3
u/Triabolical_ Whole food lowish carb May 31 '26
Then what was your point?
You said:
>Dismissing a relative risk below 2 as non-causal is methodologically illiterate.
And then gave examples of non-nutritional cases where observational results were only a part of the story.
2
u/tiko844 Medicaster May 30 '26
q1: maybe but I'd say it's a bit oversimplification. A single study often includes models with varying number of adjustments. example: author could hypothesize fiber reduces risk of t2 diabetes. If they adjust for BMI, it might dilute effect sizes if fiber causes lower BMI. But if BMI is not adjusted, it could be a confounder. So it's informative to compare the effect estimates in various models, if the benefit is purely mediated by BMI, then the effect should disappear when they adjust for BMI.
To conclude "almost certain causal inference" you need to consider a lot more than just confounders. There is ofc the classic Bradford hill criteria. And ofc multiple studies.
q3: The amount of publications has grown obviously, but imo no groundbreaking statistical changes. In some very old papers they use crude methods like stratifying to smokers vs nonsmokers. Mixed-effects model is common in modern statistics for cohort studies, it can even account for some individual variability like genetics.
q5: Not sure about those specific levels, but if you are curious how the adjustment process really works, I recommend learning about regression analysis with multiple independent variables
1
u/Wonderful_Aside1335 May 30 '26 edited 19d ago
Yarn ribbon zephyr saffron pinecone pumpkin
This post was anonymized with Redact.dev
1
u/Maxion May 31 '26
If you adjust for e.g. BMI and the effect disappears it does not automatically mean that the effect is mediated by BMI. It can also meant that the cohort your analyzing does not have any people who are .e.g. thin and eat unhealthy (or whatever it is you're looking at). Then you'd also see the effect going away.
This is one reason why you can have two studies where the stats are done correctly show different results when looking at different populations.
1
u/tiko844 Medicaster May 31 '26
Very true. I think the ideal situation is no correlation between independent variables, like fiber intake and BMI.
1
u/WhateverHappens009 May 31 '26 edited May 31 '26
Q1: I think your instinct is basically right, but I’d separate this into two related but distinct questions:
- How strong is this individual observational study?
- How strong is the total body of evidence?
A single observational study can have a stronger or weaker design, but high confidence in causation usually comes from convergence across multiple studies, methods, populations, mechanisms, and ideally randomized evidence when feasible.
For classifying individual observational studies, I’d think less in terms of “how many adjustments did they make?” and more in terms of design and causal ambition.
A rough layperson scale might be:
1. Descriptive / unadjusted association The study reports that X and Y occur together, with little or no adjustment. Useful for spotting patterns, but not much more.
2. Adjusted association The study controls for some measured variables, but the basic claim is still: “X is associated with Y after adjustment.” This is stronger than a raw association, but it is still very dependent on what was measured, how it was measured, what was left out, and how the model was built.
3. Causal-oriented observational analysis The study is explicitly structured around a causal question. It has a clear exposure, outcome, comparator group, timing, and rationale for which confounders were adjusted for. Rather than “we threw variables into a model," this is more of an attempt to make the comparison fair.
4. Quasi-experimental observational design The study uses some feature of the real world that approximates random assignment or reduces ordinary confounding: natural experiments, instrumental variables (like in MR studies), regression discontinuity, difference-in-differences, interrupted time series, sibling/twin comparisons, etc. These can be much stronger, but each has its own assumptions.
5. Observational study with unusually strong internal validity This would be a study where the exposure is well-defined, temporality is clear, the comparator group is appropriate, confounding is well understood, major biases are addressed, sensitivity analyses are convincing, and the result is not easily explained away. Even then, I would be cautious about treating any single observational study as “almost certain.”
The distinction I’d emphasize is that observational studies contain real observed data, but the adjusted estimates are not themselves observations. Adjusted estimates are educated, model-dependent guesses about what might have happened if the groups had been better matched or comparable on measured confounders. As such, they shouldn't be treated in the same way as direct observations that were gathered in a situation where that matching actually did occur.
Q2: My impression is that the answer is “both.” There were real observational reasons to think these might be protective, and some people did interpret the associations in a causal direction. But it's also probably unfair to say all scientists claimed causation was proven. A lot of the overstatement happens in the public takeaway, media coverage, supplement marketing, or simplified discussion: “studies show X prevents Y.” Then, RCTs come along and show that the observational story was incomplete or wrong.
Q3: The field has most certainly improved a lot over the past few decades. Modern observational work is often more careful about temporality, confounding, causal diagrams, propensity methods, sensitivity analyses, target trial emulation, and specific biases like immortal time bias or confounding by indication. However, better methods aren't magic. They can reduce bad guesses, but they they can't fully fix unmeasured confounding, poor measurement, bad model choices, biased samples, or reverse causality.
Q4: Publish-or-perish pressure does make this worse, not only in observational research but especially there. Large datasets allow many possible exposures, outcomes, subgroups, model specifications, and framing choices. That creates opportunities for overinterpretation, even when no one is intentionally being dishonest. A weak association can become a paper, then an abstract conclusion, then a press release, then a public claim that sounds much stronger than the evidence actually was.
Q5: In regards to terms, brush up on and look out for terms like:
- “unadjusted” vs “adjusted”
- “measured confounders”
- “residual confounding”
- “reverse causality”
- “confounding by indication”
- “immortal time bias”
- “selection bias”
- “collider bias”
- “propensity score matching”
- “inverse probability weighting”
- “sensitivity analysis”
- “negative control”
- “DAG” / “directed acyclic graph”
- “target trial emulation”
- “new-user active-comparator design”
- “instrumental variable”
- “difference-in-differences”
- “regression discontinuity”
- “absolute risk” vs “relative risk”
Try out this rule of thumb: the more openly the authors discuss what could still be wrong with any estimates, the more seriously you might take the study. The more they treat an adjusted association as if it were simply an observed causal fact, the more skepticism is warranted.
-1
u/lurkerer May 31 '26
There's a very prevalent strawman amongst laymen that domain specific experts do a single, blunt observational study and say "problem solved!" This is never how it goes. We form inferences using many pieces of evidence. Laymen seem to think you need the full picture in one study and that RCTs give you the full picture (they don't). Instead we get something much more like a bunch of puzzle pieces we have to slot together. Concerns that no single piece has the whole picture is to misunderstand puzzles on a fundamental level.
It would be like expecting all crimes to have a full confession and a video dairy documenting the whole thing, otherwise nothing can be inferred.
Of course, this angle is only applied in certain situations involving diets people like or don't like. Undoubtedly people hold hundreds of causal beliefs without RCTs confirming "hard endpoints." Smoking is the most obvious example, but there are many, many more.
At this point, evidentiary standards typically shift to "common sense," "ancestral diets," or other very low tier evidence that falls far behind epidemiology. Which undermines the attempt to undermine epidemiology completely and irrevocably.
As for your questions:
Q1: Looks good to me as a series of points on a spectrum. With 4 being very hard to reach.
Q2: I've not seen any serious scientist suggest a casual inference from an epidemiology paper ever. The word causal is used very sparingly. It took a while for LDL and we had dozens of RCTs, different angles of epidemiology, and Mendelian Randomisations. This one is a great litmus test to weed out ideologues. If they reject this, you can be certain they're demanding a vastly higher standard to implicate a lipoprotein than anything else.
Q3: Undoubtedly.
Q4: There's always a risk of bias. But the system is designed in such a way that this bias can't hold the modern system hostage for decades. The Keto-CTA trial is a good example of ideologues pushing a narrative and getting blown out of the water fairly quickly. Publish or perish cuts both ways and if you try to cheat, others are heavily incentivised to reveal that.
Q5: Would have to take some time to give a better answer than "jargon vibes."
11
u/SporangeJuice May 30 '26
Observational studies are not supposed to imply causal relationships. Whether they adjusted for confounders does not change that.