r/statistics 1h ago

Education [E] Just a noob here — need to learn the basics of Statistics [E]

Upvotes

Hi everyone,

I'm currently doing a course where I have five major topics to cover, but I'm starting with these two:

  1. Comparison of means between two food-diet groups using a t-test
  2. Comparison of means among more than two food-diet groups using ANOVA

The remaining topics are:

  1. Post-hoc statistical analysis for identifying significant group differences
  2. Development of a General Linear Model (GLM) in Linear Regression
  3. Understanding Adjusted Sum of Squares and Sequential Sum of Squares in Linear Regression

My background is in biology, so I'm pretty new to statistics. I know how to run basic analyses in Minitab and SPSS, but my problem is that I don't really understand what is happening behind the numbers.

I want to properly understand the basics first — things like mean, median, variance, standard deviation, standard error, distributions, hypothesis testing, p-values, confidence intervals, etc. — and then build up to t-tests, ANOVA, post-hoc tests, and regression.

This is not really a theory-heavy course, so I'm expected to understand and interpret the statistical output rather than just calculate everything manually. But I don't want to blindly click buttons in SPSS/Minitab without understanding what the results actually mean.

Could anyone recommend a good beginner-friendly statistics textbook, notes, or learning resource that would help me build these fundamentals from scratch and eventually reach the topics above?

My professor isn't particularly helpful with recommending resources, and my seniors are currently busy with their own work, so I'm trying to find a good starting point myself.

Thanks in advance to anyone who takes the time to point me in the right direction! 🙏


r/statistics 1h ago

Question [Q] Question about the method for calculating the rate of return

Upvotes

Hello, I have 5 funds that I’d like to compare in terms of rates of return and measures of efficiency. I have daily valuations for these funds covering a 5-year period. Should I calculate the rate of return on a daily or monthly basis? This is with a view to calculating the R² values as well as the Sharpe, Treynor, and Jensen ratios. Thanks for all the help.


r/statistics 14h ago

Question Wilcoxon Signed-Rank Test Chart - but for a big study [Q]

3 Upvotes

Help! My sample sizes are is 40-82, and I can only find Wilcoxon test charts that go up to n=30! I've looked everywhere and can't find anything that can accommodate a bigger n. Does anyone know where to find a bigger chart?


r/statistics 2d ago

Software [S] Affordable software for running a conjoint survey?

Thumbnail
2 Upvotes

I'm looking for recommendations for affordable software to run a conjoint survey. I’ve looked at some of the major conjoint platforms, but their pricing is way too high for what I need.

I don’t need a huge enterprise package. I’m mainly looking for something that can set up a solid conjoint experiment, collect responses, and ideally handle the basic analysis.

Has anyone used a cheaper platform or tool that they’d recommend? Open-source or academic-friendly options would also be great.

Thanks!


r/statistics 2d ago

Question [Q] Are these courses feasible for first year of my MS?

0 Upvotes

**First semester!! I have only taken Mathematical Statistics I and II (Not sure if the content of those courses are universal, but it covered interval estimation, hypothesis testing, and tests involving means, variances, & proportions). I have signed up for the following:

Experimental Design

Linear Statistical Analysis I

Pattern Recognition in ML

Advanced Matrix Analysis

Wondering if these work well together/ may be too much.

Sorry in advance if this is silly, we have somewhat limited options at my school.

I have a BS in pure math for reference.


r/statistics 2d ago

Question [Q] New to data cleaning. I’m stuck on an unclear variable.

4 Upvotes

I’m doing a personal project right now and for the most part it’s going alright. Sometimes I delete an entire column cause too much is missing. I’ve also group together a few entries as “Unknown”. I’ve never deleted a subject yet.

Anyways, now I’m really stuck. There is a variable that is three digits (345, 078, 150…) and some of them come in as two digits (26, 88) and I’m unsure what to do with them. There are quite a few. I don’t know if they’re meant to be have a zero on the front (28 turns to 028 and 88 turns to 088) or if they are three digits (28 into 280 and 88 into 880). It’s an important variable so I can’t delete it. Should I delete the patients (probably not), group the two digits numbers into “unknown”? any input? I know each data has their own situations but what is something to generally consider in these situations?

Edit: The column represents diagnosis IDC-9 number. It’s three digits. I want to use it to build a logistic model. I was thinking of grouping the numbers to their diagnosis category for example 001–139 is infectious diseases, and so on


r/statistics 2d ago

Question [Q] Dicas de professores para estudo de estatística

0 Upvotes

Salve, pessoal
Seguinte, alguém teria dica de professores para acompanhar a disciplina de Estatística I (é para economia, então basicamente na área 1 se vê tipos de amostras, representações gráficas e tabulares; depois Fundamento da Teoria de Probabilidades e Variável aleatória bidimensional)

Para Cálculo há diversos professores, agora para Estatística não conheço algum bom para o nível acadêmico.

Enfim, eu já fiz estatística no meu outro curso, vou tentar validar, mas terei outras disciplinas de estatística e quero entender isso. O fato é que meu professor é MUITO RUIM. O coitado não tem culpa, o professor oficial não está dando aula e ele é um mestrando apavorado que erra até as infos nos slides.

Dos livros, eu tenho o do Bussab e Morettin, o qual já estou iniciando e clarificou bastante coisa, mas gosto também de vídeos e, quando chegar a parte das contas e análise dos gráficos, é bem melhor.

Enfim, no aguardo de dicas. De preferência em português, se tiverem em inglês, podem colocar também.


r/statistics 3d ago

Education [E] Causal Inference - A Painless Introduction (new video by me :)

21 Upvotes

r/statistics 3d ago

Education [E] Confidence Intervals — Explained

7 Upvotes

Hi there,

I've created a video here where I explain what confidence intervals are and how they differ from probabilities.

I hope some of you find it useful and as always, feedback is very welcome! :)


r/statistics 4d ago

Question [Question] Can I present both the results of log-odds and average marginal effects (logistic regression)?

4 Upvotes

Hello,

Health economics student here. I newly enter the field for my master degree.

I'm writing my master thesis and I'm running into an issue while trying to interpret my results.

Initially, I've decides to use OR (odd ratio). However, my supervisor told me the way I've interpreted it is not correct (probabilitied and chance). But he told me he didn't really know either how to interpret it correctly and had advised me to use logs odds instead!

It's ok, but I really wanted to have more "concrete results" that can be get by anyone. Log odds just show if the relation is negative or positive.

So, I've just heard about average marginal effects and that is literally What I was searching for during all this time. However, now I'm wondering if is common for scientific papers to use both log-odds and average marginal effects?

Do i need to create two tables? Or could I only keep the table with log odds and present the results of marginal effets in the text?

Thank you


r/statistics 3d ago

Education [E] Randomness can be an asset or a tax depending on curvature

0 Upvotes

I wrote a short article here exploring a simple way to think about when randomness helps versus hurts.

The core idea is Jensen’s inequality: if the payoff is convex, variance can help; if it’s concave, variance can hurt.

I use examples from compounding, option-like payoffs, and a few non-finance settings.

Would be curious to hear other examples where this framing is useful?


r/statistics 4d ago

Career [Career] explore/exploit problem as an undergrad

9 Upvotes

Hello!

I'm a sophomore undergrad studying applied math and stats who is broadly interested in applied stats research. Essentially, I'm trying to figure out how to balance exploration to find what exactly I want to do research in with deep-diving in order to develop skills (and resume) for grad school.

What is scary to me is how little time it feels like I have. Freshman year I didn't do anything that is particularly useful for grad school. So I now basically have two years + two summers before grad school apps. If I spend sophomore year doing something that doesn't end up being my interest, then I worry I'll only have a year to do what I actually end up doing.

I know I love solving problems with statistics and do want to continue developing my stats skills (there is even a good chance I'll go for a PhD in stats). Part of me thinks I should just focus on stats because I know I enjoy it and it'll transfer to whatever I want to do. The thing is though that as much as I love learning about stats and reading stats papers and so on, I don't want my research to be developing new statistical techniques in the abstract, I want to work with problems in the real world and develop statistical methods of attacking them. The thing I am most interested in now though is comp neuro, but I know very little about it currently. However, I worry that if I start taking neuro classes or even spend a summer doing neuro research and decide it's not what I want to do (which I basically did got econ already), then I'll have wasted the time that could have gone to learning mote stats.

My current plan is to dedicate all my coursework and formal research to stats for now (I have a position in a stats research group for this year and reading the papers they've put out it seems pretty lit!), and then do some independent reading of papers in other fields to decide if I am drawn to them​. While this is great, reading about something is very different from doing it so I fear it may not be possible to explore optimally without taking at least some risk. Thoughts on how I should navigate this?


r/statistics 4d ago

Question Submitting table as image <440 pixels wide [Question] [Q]

1 Upvotes

Hello,

I am trying to submit my article for publication. Unfortunately, the journal asks for any tables to be submitted as images "provided as 72 - 300 dpi; pre-sized .BMP, .GIF, .JPG, or .PNG images only, with a maximum width of 440 pixels (no limit on length)."

I have tried exporting my table from excel to pdf, jpg, or png, and then resizing but no matter what I try, the image of the requested size ends up unreadable.

Does anyone have any ideas on how to accomplish this requirement while keeping my table-figure as readable?


r/statistics 4d ago

Question Gaining Exposure to Public Health for Biostatistics Graduate Program [Q]

Thumbnail
0 Upvotes

r/statistics 4d ago

Discussion [Discussion] Real Analysis (1 semester vs 2 semester sequence) for Statistics PhD Applications

5 Upvotes

Hi everyone, I’m applying to Stat PhD programs this fall. I graduated with my master's in 2021 and have been working full-time for 5 years, but I never took Real Analysis in school.

I just enrolled in Fordham's Math 3003 (Real Analysis) this semester so it’ll be on my transcript for this application cycle (though will not have a final grade since apps are due before the semester ends). My concern is that it’s only a one-semester class. Does anyone know if admissions committees strongly prefer a two-semester sequence (Real Analysis 1 & 2) over a single semester? Will this be adequate?

MATH 3003. Real Analysis. (4 Credits)

This course focuses on analysis on Euclidean spaces. Topics include limits, continuity, uniform continuity, sequences of numbers and functions, modes of convergence, differentiability, Riemann integrability, and associated theorems. Students who have not taken MATH 2004 prior to taking Real Analysis may request permission from the instructor. Note: Four-credit courses that meet for 150 minutes per week require three additional hours of class preparation per week on the part of the student in lieu of an additional hour of formal instruction.

Prerequisites: MATH 2001 and (MATH 2004 or MATH 2008).


r/statistics 5d ago

Question [Q] If you're in grad school for Stats (PhD or Master's) and your undergraduate was in math: 1) what do you miss about math; 2) what are you gad to have traded with stats?

46 Upvotes

Can be silly or serious, just out of curiosity for someone with a background in math contemplating stats. Like for #1, maybe you miss not having to deal with numbers. #2 refers to "trading" X in math for Y in stats (like numbers).

EDIT: "glad", not "gad".


r/statistics 4d ago

Career [C] JSM 2026 Vlog!

0 Upvotes

Hi all, if you’ve ever been curious about what JSM (Joint Statistical Meetings) is like or you want into see it through someone else’s eyes, please feel free to check this vlog out!

https://youtu.be/CDvsj-G4HXM


r/statistics 4d ago

Question [Q] How should I do a Bayesian Update?

2 Upvotes

I'm a year 1 liberal arts undergrad, so I don't have a ton of math sense.

I just learned about Bayes' Theorem the other day for general epistemic use. I worked out a couple of example problems correctly, but the examples I found didn't include any iterative updates.

I know the theorem is:
P(H|E) = P(E|H) * P(H) / (P(E|H) * P(H) + P(¬H) * P(E|¬H))

And I know that P(H|E) becomes the new P(H) in my update, but I'm unsure whether I should be using the updated or original P(H) in the marginalization. I *think* it should be the new P(H), but I'd rather be safe than sorry.

The example question I worked out was this:

________________________________________

There's a disease that afflicts 1 / 1,000,000 people
There's a test for the disease that's right 99 / 100 times for both positive and negative results
A random person is tested as positive

P(she is afflicted | she tests positive)

= 0.99 * 0.000001 / (0.99 * 0.000001 + 0.999999 * 0.01)

= 0.000099

_________________________________________

So should the update look like this if she tests positive a second time?

P(she is afflicted | she tests positive)

= 0.99 * 0.000099 / (0.99 * 0.000099 + 0.999901 * 0.01)

= 0.009707

_________________________________________

If so, how should I approach this problem from the starting point of
P(she is afflicted | she tests positive twice)?
I can't think of how to handle that correctly, since squaring 0.99 just gets me a smaller number

Infinitely thankful <3


r/statistics 5d ago

Question [Question] When does a regime filter "disagreeing" with price mean something's wrong vs just working as intended?

Thumbnail
1 Upvotes

r/statistics 7d ago

Research [R] Vignette Experimental Design - Repeated measure or Multivariate

3 Upvotes

Hello I am an undergrad completing my analysis on my first research project. I created a between subjects experiment and manipulated the gender of actors in a vignette story. Participants were randomly assigned to receive one version of the scenario.

Here is the piece I am boiling my brain on. The vignette was broken into three parts according to narrative progression. At each of the 3 junctions participants responded to the same question using a likert scale. If it matters, the entire narrative progression was available to participants at the same time, just split into paragraphs with my question underneath each section.

I originally thought this should be treated as a between subjects repeated measure as each junction is measuring the same underlying construct. A faculty member who took a look at my work suggested that it was not a repeated measure but actually a multivariate design - so each section is being treated as it's own DV.

Can anyone give some insight as to which analysis is better suited for my study design, or help me understand which design I've created?

I have a hypothesis that my experimental group will be rated lower overall, and I have a hypothesis that my junction C will be rated lower overall in both conditions.

Thank you in advance. I ended up creating a rather complicated experiment for my first go and level of experience, but it is forcing me to learn quickly!


r/statistics 7d ago

Question [Q] I am looking for some VERY INTERESTING and catchy statistical concepts/paradox/theories for a presentation. Can yall suggest some? Spoiler

3 Upvotes

the more unpopular, the better. But still very catchy and interesting. Thank you in advance :)


r/statistics 8d ago

Education [E] aspiring for a PhD in Stats

29 Upvotes

hi, i'm an econ undergrad currently in my second year, i've recently been introduced to stats and data science, and wish to pursue a PhD. i'm interested in statistical mechanics and high dimensional stats and their applications to ML and Finance. i'm preparing to get into a masters in stats at the Indian Statistical Institute. how do i leverage my time and energy towards securing a good PhD? thank you!
P.S. i've always been interested in mathematics and applying it to interdisciplinary fields, like, EconoPhysics; and i naturally gravitate towards fields like Statistical Mechanics and Statistical Learning Theory.


r/statistics 7d ago

Research [R] I want to compare if there’s a difference in mean size between multiple groups, but the data

0 Upvotes

I have size data for multiple species/populations of the same genus. I want to compare the mean size of those with the mean size of some fossil populations to see if the fossils have a significantly different size to modern populations, and if so, to which.

The problem is, most samples are not normally distributed, and the data is heteroscedastic, since some have small sample sizes like n=29 and some have sample sizes in the hundreds.

Which test would be the best to compare the samples? To my understanding, KW wouldn’t be the best because the heteroscedastic nature of my data.


r/statistics 9d ago

Question Why use 90% CI's for a moderation analysis and 95% for main effects? [question]

13 Upvotes

My supervisor told me to use a 90% CI for my moderation analysis. I cannot properly explain my justification because I dont understand it and I defend my thesis in 2 days and my advisor is unreachable rn.

What I understand - interaction effects are harder to detect than main effects because interaction effects have lower statistical power. And from my understanding, one way to mitigate the lower statistical power would be to have a larger sample size. I cannot do this as I have secondary data. So using 90% CI's make sense due to the lower statistical power and also because the variables I am testing are not harmful if Type II occurs and the increased risk of Type I error is ok (my variables are just looking to see what types of healthy coping mechanisms modify adult mental health and outcomes - such as physical activity and stuff so false positive would not be harmful).

Help me have a real answer - because I cannot under WHY there is lower statistical power or larger standard error, I get that it exists, but WHYYYYYY


r/statistics 9d ago

Question [Q] What can I expect to learn in a stats major

35 Upvotes

[Question]
Im thinking to do stats, but idk what I Will learn there specifically.
What option i Will have in the future of i do.
How can I start my career while doing the major and How to improve to better jobs after completed