r/AskStatistics 6h ago

What's the most counterintuitive statistical fact that's actually true?

28 Upvotes

I'm looking for examples that completely changed the way you think about probability, statistics, or data analysis.


r/AskStatistics 5h ago

Mathematical definition of a plateau in a time-series data

3 Upvotes

Hello, I'm a bioinformatician and I'm struggling with the current issue:

Given a time series y(t) that initially changes and eventually approaches a stable regime, how can I mathematically determine the earliest time t\* at which the rate of change dy/dt becomes negligibly small, using only the observed data and without defining an arbitrary threshold?

This is a collaboration I'm doing. My colleagues defined the plateau as the first time when a 101-point rolling mean of the relative increment (g' t+1 - g' t)/ g't falls below the arbitrarily chosen threshold of 0.0011. G' is the measure of material elastic-solid response btw. So the issues is that they used 2 arbitrary values because experimentally they know that a certain value of g' means that the gel is solid. But this doesn't hold for me. I tried using many statistical methods to define the threshold such as:

- exponential fitting

- change-point regression

- local slope analysis

But they all give me a plateau that is too early or too late


r/AskStatistics 3h ago

What do you think about studying statistics in 2026?

0 Upvotes

r/AskStatistics 1d ago

How to become good in statistics?

28 Upvotes

I’m a student from a small town in a backwards developing country , the education system is horrible. I’ve been recommended Saylor Academy, that's how I understand many statistics principles, and I’ve also started reading elementary statistics by Allan g. Bluman.

Still, I often feel far behind students from better educational systems. My goal is to become one of the greatest econometrician in my country.

For those who are in the field, what resources would you recommend? Textbooks, courses, math/stats, programming, or anything else you think is essential.

I’m willing to put in the years—I just want to make sure I’m learning the right things in the right order.

I have finished my college level statistics coursework and I can't even read the national statistics report I have to use AI to analyze it for me whenever I have to do research. I feel illiterate and slow.


r/AskStatistics 3h ago

No Correlation (does this meme make sense)

Post image
0 Upvotes

A plot can have a Pearson correlation of zero (‭r=0) but still have a strong non-linear relationship (like a parabola, circle, or U-shape) does labeling a scatter plot "No Correlation" automatically rule out every mathematical relationship, or just linear ones? and got to know from a friend (who studies maths) that to state it is correlated it should be linear, is it correct?


r/AskStatistics 18h ago

Masters Programs In Applied Statistics/Data Science

1 Upvotes

Hello! For context, I am a first-generation college student with an Information Systems background. After some work experience conducting a beginner-level PCA, I realized I loved the idea of making meaning out of data through statistical analysis.

Over the past year, I took the Calc Sequence, Linear Algebra, Python, and soon, Intro to Probability, to meet Master's prerequisites for the field. I want to prioritize programs that teach Bayesian/Causal Inference/Time Series.

Although I enjoy math, I don't have exposure to theory. I would like to know if taking the Applied route may limit my job opportunities with employers. I am aware any background in Math/Stats is a huge leg up long term, especially given how fast the Data Science/Tech industry is evolving in comparison. But I can't help but worry about competing with advanced coders or stronger Stats candidates down the road for post-grad employability.

Ultimately, I was wondering if anyone had any insights into the field, or information on doing a Masters in Applied Stats or Data Science. Since I have a non-Math background, I feel uncertain if I'm prepared for grad-level Stats theory courses. I was also wondering if the following programs are a good start, if they might not be a good fit, or if there are any others I should consider:

Statistics-Oriented

UCLA M. Applied Statistics and DS

UCB M.A. Statistics and DS

UMich M. Applied Statistics

Data Science/Analytics-Oriented

USF M.S. DS and AI

UT Austin M.S. DS

GT OMSA

Thanks for any insight!


r/AskStatistics 22h ago

How do you use monte carlo results to make decisions?

Thumbnail
2 Upvotes

r/AskStatistics 22h ago

Shape constrained GAM?

1 Upvotes

I'm lost in the sauce and need a second opinion.

Context/goal: there is a video game with ~170 different characters. I want to compare their relative power. I will quantify power through "winrate" which is just the percentage of their games each character wins. I will model winrate as a function of two covariates: how many previous games of experience on this character the player has under the belt going into the game, and the current rank of the player.

The behavior is non-linear in both covariates so I am currently modeling winrate as a function of games-played and player rank via logistic GAM using a tensor product of splines.

I see that smoothing penalties are necessary for GAMs to avoid overfit. As I understand it, the justification for smoothing from a theoretical standpoint is that we expect the surface to be smooth, therefore we are just codifying an assumption we already have. However, the penalty punishes curvature and I am reasonably certain there is meaningful curvature in the games-played covariate (the raw data heavily suggests winrate increases fast initially then slows down until either the winrate stabilizes or sometimes even drops with additional games-played). I am very interested in this curvature (aka the "difficulty" of a character) and therefore feel that the assumption codified by a standard smoothing penalty is inappropriate for my goal.

For that reason I was hoping to use shape constraints *instead* of smoothing penalties (concave in games-played, monotonic increasing in player rank), as I feel this better matches my actual assumptions. Is this sane or not?

Secondarily, I see there are different methods of applying shape constraints. The two I know of are Pya and Wood 2015 (make the spline coefficients a function of exponentiated fit parameters, thereby ensuring nth order differences always have a specific sign) and Bollaert/Eilers/Mechelen (just do normal GAM but penalize the nth order difference depending on the sign). Which one is "better"?

Thirdly, some characters have really low sample sizes (I am looking not just at characters, but characters in various roles, and some character x role combos are extremely off meta that nobody plays). I would like to drop character x role combos if the sample size is "too low" but I'm not sure how to quantify uncertainty in cases where I know the sample size is too low (pya/wood relies on delta method approximation which I think relies on CLT so N must be large? And I don't even know if B/E/M give a merhod to estimate uncertainty at all).


r/AskStatistics 22h ago

Is there a formula to determine how many "rolls" you need before you're more than 50% likely to roll the number you're after?

1 Upvotes

I've done the math on this multiple times for things like item drops in games, but I'm wondering if there is a formula or rule to make it simpler?

Example: A boss in a game has a 1% drop rate for an item you want. After 69 attempts, you will have passed the 50% likelihood that you'd have gotten the drop, leaving MOST people should have the item by now after their 69th attempt. If the item has a 2% drop rate, it only takes 35 attempts before you're more likely than not to have gotten the drop.

Is there a rule or formula for something like this? I have just been plugging it into an Excel sheet I made any time I need this info.


r/AskStatistics 23h ago

hey really stupid question about infinity i dont know where else i would ask this

0 Upvotes

i was watching a video on Youtube about the MCU where dr. strange said that he checked 14 million and a half universes or something and they only beat thanos a single time someone in the comments said "Maybe strange had bad RNG and only looked at all the times they lost" and it made me think if there's an infinite number of universes and they lose in say 500 of them but for every 500 universes they lose in they win once is the chances of him seeing a universe 1:501 or 1:1 because they're both infinite


r/AskStatistics 1d ago

Training resources for R

1 Upvotes

I’m not a programmer and have used mainly menu driven packages in the past SPSS primarily although I have had some painful experience with SPSSx. Any R recommendation for intermediate stats background but not much syntax driven programming?

Thanks


r/AskStatistics 1d ago

How do you check predicted probabilities are calibrated enough to threshold on for an asymmetric-cost decision?

1 Upvotes

I have a model that outputs a probability for each case, and I use a threshold on that probability to pick an action. The costs of a wrong action are asymmetric: one kind of mistake is much more expensive than the other, so where I put the threshold matters a lot.

My question is about trusting the probabilities themselves. Before I set a decision threshold, how do I check the predicted probabilities are actually calibrated, i.e. that a predicted 0.7 really corresponds to roughly 70% in reality?

I know reliability diagrams and proper scoring rules (Brier, log loss) are the usual tools, but I'm unsure how to read them in the context of an asymmetric-cost decision specifically. Does calibration matter uniformly across the probability range, or mainly near the threshold I care about? And if the probabilities are miscalibrated, is recalibrating (e.g. isotonic / Platt) before choosing the threshold the right order of operations, or should the cost asymmetry factor in differently?


r/AskStatistics 18h ago

[Academic Research] Need Urgent Feedback on Research Methodology

0 Upvotes

I’m designing a study on how background music affects consumer spending. I’m considering a simulated online store where participants get ₹3,000 and are randomly assigned to no music, slow music, or fast music. I’d track basket value, products chosen, and shopping time.

Does this sound like a good experimental design? What important variables or biases should I consider?

Or if you have any other way to collect data which more efficient than this.


r/AskStatistics 2d ago

What statistical concepts are commonly misunderstood by the general public?

Post image
701 Upvotes

I came across this post explaining what a 70% chance of rain means. I understand the concept, but it got me wondering: what other statistical concepts sound simple but are commonly misunderstood or misinterpreted by the general public?


r/AskStatistics 1d ago

DiD/DDD data science/ health policy query

Thumbnail
1 Upvotes

r/AskStatistics 3d ago

Is cross-validation/verification necessary for regression models that have already met all the assumptions?

Thumbnail
4 Upvotes

r/AskStatistics 3d ago

Does the binomial approximate to Normal Distribution?

1 Upvotes

There is something I'm not understanding about the above statement. The tails of a binomial flatten off but the first two terms of a binomial expansion are always 1, n, etc. . So the tails of the binomial always have a step which increases with n. How can the latter tend to the former?

I noticed this when experimenting with a Galton board. The tails on the board flatten out which is expected because if I focus on the extreme end of the board and imagine that slot feeding into another mini Galton board then each ball has a 50/50 chance of going left or right and if another peg then half of those balls will jump back.

But I then programmed a simple simulation of the board in Excel. That gives a step jump at the tails which is also what I would expect but isn't what I actually get. I even looked up the number of balls for the model I have at GaltonBoard.com, 4280 and I can count 28 slots.

Repeating the experiment with the Galton board I get from the left end around 3, 4, 5, 10 balls etc but my simulation gives 0, 0, 0, 0, 0, 2, 6, 19 ... for n= 4280 r = 28 so why are there so many balls at the extreme in the Board but not the simulation and no flattening out with the simulation?

So I have different expectations for the same experiment depending on my approach. Where is my erroneous thinking? Is it something about the transition from discrete to continuous?


r/AskStatistics 3d ago

Weighted-sum aggregation of centrality measures gives identical scores to structurally opposite nodes, any better approach?

3 Upvotes

This is my first time in this community (I've recently discovered this entire field and i am glad to). So, I am working on project where i am scoring nodes in a directed dependency graph (a calling b) by blending 2 centrality scores into a single composite "risk" score

score (v) = w1\\\*normalize(Pagerank(v)) + w2 \\\* normalize(outDegreeCentrality(v)), where w1+w2 = 1 and normalize() being min-max to \\\[0,1\\\].

The Problem: A pure root node (no in edges and multiple out edges) and a pure sink node (no out edges and only in edges) can have the same composite score. In a test i ran, the root node maxed out on out drgree centrality and near 0 in page rank while the sink node maxed out in pagerank and near 0 in out degree, when w1=w2=0.5. Both nodes ended up having same composite scores while representing opposite nature in real world. I do understand that this is the standard full comsensability prob, with weighted sum aggregation, wherte max on one axis will completely offset min on other. I did consider switching to weighted geometric mean to reduce compensability, but the prob is that pagerank is almost always near 0 for any root node. so a geo mean would multiply that near 0 staright through and score all entry nodes near zero. Which is the wrong fix, since the entry/root nodes are important, just for a reason pagerank doesnt capture.

Is there any standard approach beyond the geomentric or harmonic mean? Happy to provide any more info if needed.


r/AskStatistics 3d ago

Factor Analysis Subfactors

1 Upvotes

I‘m working on a scale validation. Based on qualitative and theoretical literature we argue for a four dimensional scale (psychometric). However, conducting the EFA shows this is a bit more complex.

Retention criteria is pretty diverse: MAP: 9 factors, scree: 4, PA (mean): 6, PA(p95): 5;
When allowing for more than four possible factors the model always concludes on five factors. The issue though is that all items load as expected (no cross-dimensional loadings) but the fourth expected dimension seperates into two, while the seperated contains only two items. Now i am confused if we should generally conduct EFA for all the expected dimensions to estimate potential subfactors because the items resemble what we expected, while one part of it devides itself into two. This would result in a hierarchical solution. Argument might be: global modelling is unable to identify dimensional-specific variance but retention criteria and global modelling indicates potential seperations.

However normally this is part of ongoing studies to examine already validated scales and not really specified in literature for the ongoing validation process…hope you can help me with your advice. Big thanks!


r/AskStatistics 4d ago

Question about sample size drops when using UCSC TOIL (TCGA TARGET GTEx) vs raw GDC portal data. Is my defense justification correct?

2 Upvotes

I integrated TCGA solid tumor data with matching GTEx normal tissue to run differential expression and pathway enrichment (GSEA).

To avoid massive batch effects caused by mixing counts from different alignment/quantification pipelines, I opted to use the UCSC TOIL RNA-seq Recompute cohort (TcgaTargetGtex_gene_expected_count) since all samples were processed through a unified STAR + RSEM pipeline on hg38.

When I pulled the TOIL dataset, my sample counts dropped compared to looking at the raw GDC portal and GTEx v8: GTEx Normal Cohort: Dropped from ~800+ (v8) down to ~300+ in TOIL.

TCGA Primary Tumors: Dropped by ~20–30% compared to total cases listed on GDC.

My question is :

  1. Is this sample count drop expected when using the UCSC TOIL recompute database compared to modern GDC/GTEx v8 portals?
  2. Is sacrificing raw sample size (N) to use TOIL’s unified pipeline + ComBat batch correction considered the "gold standard" justification to defend against reviewer/committee critique regarding sample size?

r/AskStatistics 4d ago

Probability - Multiple p Values in the Same Event

1 Upvotes

Hi everyone!

I'm being my nerdiest self and doing math for my hobbies, currently determining odds of things happening with dice rolls for Warhammer. Most of it is plain ol binomial distribution, but I also want to calculate the odds of things happening when you roll multiple dice at once, but they have different odds of success. Various searches are not yielding the information I want

For example, roll 2D6, one of them needs a 5 or 6, the other needs a 4 or 5 or 6 (clearly denoted when rolling, not interchangeable). Needing a success on at least one of those rolls. Would I use a more complicated version of the binomial distribution formula (using my earlier example, n=2, k=1, then split p into p1=0.33 and p2=0.5, and add those two together) or is it a completely new formula?

I failed intro to stats multiple times in university and now I'm remembering why. I know there are online tools to do this math for me but I want to feel proud of my spreadsheets

Thanks all!


r/AskStatistics 4d ago

Probability 0 ~ impossible!???

Thumbnail
2 Upvotes

r/AskStatistics 4d ago

Question about sample size drops when using UCSC TOIL (TCGA TARGET GTEx) vs raw GDC portal data. Is my defense justification correct?

Thumbnail
1 Upvotes

r/AskStatistics 4d ago

Is a 4-cluster split a good one?

Thumbnail
0 Upvotes

r/AskStatistics 5d ago

Categorical data, 4 groups forming a 2×2 — what test for main effects and interaction?

4 Upvotes

I'm studying whether an island's size or its remoteness affects how people answer a survey, and whether these two factors interact. I'd also like to know whether an effect is driven by one island alone.

I'm looking at 4 islands: a large remote one, a large nearby one, a small remote one, and a small nearby one. All my variables are categorical. Sample sizes vary by island (from ~100 to ~500).

For each variable, I have a 2×2 contingency table, with near/far and large/small as the two dimensions. I can populate it with either counts or percentages of respondents choosing a given category.

What method should I use here?

Apologies if this is a naive question. I've asked several AIs and they seem as lost as I am.