Blog
Plain-English guides to choosing and reporting statistical tests, and how we build AnalyEase.
Featured
How we validated AnalyEase's statistical engine against SciPy and pingouin
Every test in AnalyEase is checked against established reference implementations, to four decimal places on p-values. Here is how the validation works, what it caught, and the limits we tell you about instead of hiding them.
Statistical rigour, 7 min read
Choosing a test
6 min read
When two groups become three, running more t-tests quietly inflates your false-positive rate. How t-tests, one-way, two-way and repeated-measures ANOVA fit together.
Choosing a test
6 min read
What "parametric" actually assumes, the non-parametric equivalent of each common test, and why normality tests can mislead with very small or very large samples.
Choosing a test
5 min read
Before-and-after measurements, matched samples and independent groups. How to tell which design you have, and the pseudoreplication mistake to avoid.
Multiple comparisons
6 min read
Your ANOVA is significant, so which groups differ? What each correction controls, and how to pick between all-pairs and versus-control comparisons.
Figures
4 min read
Standard deviation and standard error answer different questions. What each tells a reader, why SEM bars shrink as n grows, and what to write in your legend.
Reporting
4 min read
Power calculated after the experiment is a direct function of the p-value, so it adds no new information. Report effect sizes with confidence intervals, and plan power in advance.
All posts
Statistical rigour, 7 min read
How we validated AnalyEase's statistical engine against SciPy and pingouin
A figure is only as trustworthy as the numbers behind it. Before any statistical test appears in AnalyEase, it has to reproduce the output of an established reference implementation on a fixed dataset.
Why validation matters
A wrong p-value doesn't announce itself. It looks exactly like a right one, sits in a figure legend, and can travel all the way into a published paper. So we treat the statistical engine the way a lab treats a new assay: nothing is trusted until it has been checked against a known standard.
Checked against reference implementations
Each test is validated against output from widely used, peer-scrutinised statistical libraries: SciPy for core tests such as t-tests, Mann-Whitney U, one-way ANOVA, Kruskal-Wallis and correlation, and pingouin for the ANOVA family, including repeated-measures and mixed designs. As new analyses are added, each one names its reference tool before it is built.
Locked fixtures with tight tolerances
For every test we keep fixed datasets, called fixtures, whose expected results were computed once by the reference tool and then locked. AnalyEase must match them within:
- ±0.0001 on p-values and correlation coefficients
- ±0.01 on test statistics
Those tolerances allow for normal floating-point differences between implementations, not for real disagreement about the answer.
Checking the checkers
Reference values can be wrong too, for example when a number is copied from the wrong dataset. Locked values are cross-checked, and any that look doubtful are re-derived independently, by a second reference run and, where practical, by hand. When an expected value didn't reproduce, we traced where it came from and corrected the fixture. We never adjusted the engine to match a number we couldn't explain.
Tests before code
New statistical features are written test-first: the check is shown to fail before the feature exists, then pass once it is built. The complete automated test suite runs before any change is accepted, because a change in one test can quietly break another.
Real data sizes, not just toy examples
Small textbook datasets hide problems that appear at real scale. Features are also checked with realistic experiments, such as timecourses with dozens of timepoints and many-group designs, and the resulting charts and tables are reviewed by eye, not only by an automated assertion.
Limits we show you
Rigour also means being honest about what a test can and can't do. With very small samples, rank-based tests have a hard ceiling on how small their p-value can get, however large the true difference. AnalyEase shows caveats like this next to the result rather than hiding them, and it never reports post hoc power, a statistic that adds nothing to the p-value.
Deterministic by design
The engine uses fixed rules, not AI. Give it the same data and it returns the same answer every time, which is exactly what makes validation against a reference possible in the first place.
Join the private beta
All posts
Choosing a test, 6 min read
T-test vs ANOVA: which test should you use to compare means?
Both compare group means. The difference is how many groups you have and how your experiment is structured.
Two groups: the t-test
A t-test asks whether the means of two groups differ by more than you'd expect from random variation. Control vs treated, wild type vs knockout: if there are exactly two groups, a t-test is the standard choice.
Three or more groups: one-way ANOVA
Once you have three or more groups, such as vehicle and two doses, use a one-way analysis of variance. ANOVA tests a single question first: is there any difference among the group means?
Why not just run several t-tests?
Each t-test carries its own 5% chance of a false positive. With three groups there are three pairwise comparisons, and the chance of at least one false positive rises to about 14% (1 − 0.95³). With six groups there are fifteen comparisons, and that chance exceeds 50%. ANOVA followed by a corrected post hoc test keeps the overall error rate where you set it.
Two independent variables: two-way ANOVA
If your design crosses two factors, such as genotype and treatment, a two-way ANOVA tests both factors and their interaction: whether the effect of treatment depends on genotype. Interactions are often the most interesting result, and separate one-way analyses can't detect them.
The same subjects measured repeatedly
When each animal, patient or culture is measured more than once, for example at several timepoints, the measurements are not independent. Use a paired t-test for two timepoints and a repeated-measures ANOVA for more. Designs that mix between-group and within-subject factors need a mixed-design ANOVA.
Assumptions to check
- Independence: observations in different groups don't influence each other.
- Approximate normality of the data within each group, or of the residuals.
- Similar variances across groups, or a version of the test that doesn't assume them, such as Welch's.
If normality clearly fails, a non-parametric equivalent is usually the better choice.
Quick summary
- Two independent groups: unpaired t-test
- Two measurements on the same subjects: paired t-test
- Three or more independent groups: one-way ANOVA with a post hoc test
- Two crossed factors: two-way ANOVA
- Repeated measurements over time: repeated-measures or mixed-design ANOVA
Join the private beta
All posts
Choosing a test, 6 min read
Parametric vs non-parametric tests: how to choose, with examples
Parametric tests are more powerful when their assumptions hold. Non-parametric tests are safer when they don't.
What "parametric" means
Parametric tests, such as the t-test and ANOVA, assume your data follow a particular distribution, usually a normal one, and compare means. Non-parametric tests make fewer assumptions. Most work on the ranks of the values rather than the values themselves, which makes them robust to skew and outliers.
Common tests and their non-parametric equivalents
| Design | Parametric | Non-parametric |
| Two independent groups | Unpaired t-test | Mann-Whitney U test |
| Two paired measurements | Paired t-test | Wilcoxon signed-rank test |
| Three or more independent groups | One-way ANOVA | Kruskal-Wallis test |
| Repeated measurements | Repeated-measures ANOVA | Friedman test |
| Association between two variables | Pearson correlation | Spearman correlation |
Checking normality, and its pitfalls
The Shapiro-Wilk test is a common formal check. It has two blind spots worth knowing:
- Very small samples: with three or four values per group, it has almost no power, so "passed" doesn't mean the data are normal.
- Very large samples: it flags trivial departures from normality that wouldn't affect a t-test at all.
Always look at the data too. A dot plot or histogram often tells you more than a single p-value.
The trade-off
When data really are normal, parametric tests detect real effects with fewer samples. Non-parametric tests give up some of that power in exchange for robustness. They also have limits with tiny samples: with three values per group, the smallest possible two-sided p-value from a Mann-Whitney test is 0.1, however different the groups are.
Other options for skewed data
For right-skewed measurements such as concentrations or tumour volumes, a log transformation often makes the data close to normal, letting you use a parametric test on the transformed values. Report clearly that you did so.
Join the private beta
All posts
Choosing a test, 5 min read
Paired vs unpaired t-test: what's the difference and when to use each
The choice depends on one question: is each value in one group linked to a specific value in the other?
Unpaired (independent) t-test
Use an unpaired t-test when the two groups contain different, unrelated subjects. For example, tumour weight in eight vehicle-treated mice compared with eight different mice given a drug.
Paired t-test
Use a paired t-test when each measurement in one group has a partner in the other:
- the same cells or animals measured before and after treatment
- left and right sides of the same subject
- samples split and processed under two conditions
The test works on the differences within each pair, which removes variation between subjects. When that variation is large, a paired design can detect an effect that an unpaired analysis of the same data would miss.
Welch's or Student's?
For unpaired data there are two versions. Student's t-test assumes both groups have similar variances. Welch's t-test does not, and it performs well even when variances are equal, which is why many statisticians recommend it as a sensible default.
The mistake to avoid: pseudoreplication
Technical replicates, such as three wells from the same culture or several fields from one section, are not independent samples. Treating them as separate data points inflates n and makes p-values look far more convincing than they are. Average technical replicates first, then analyse one value per biological replicate.
Quick check
- Could you draw a line connecting each value in group A to one value in group B? Use a paired test.
- Are the groups made of entirely different subjects? Use an unpaired test.
Join the private beta
All posts
Multiple comparisons, 6 min read
What is a post hoc test? Tukey, Dunnett, Holm-Šidák and FDR explained
A significant ANOVA tells you that some groups differ. A post hoc test tells you which ones, while keeping false positives under control.
Why corrections are needed
Every extra comparison is another chance of a false positive. Post hoc procedures adjust p-values, or the threshold they are compared against, for the number of comparisons in the family.
Tukey's test: every pair
Tukey's honestly significant difference test compares every group with every other group. Use it when all pairwise differences are genuinely of interest.
Dunnett's test: each group against a control
If your question is only "does each treatment differ from vehicle?", Dunnett's test compares each group with a single control. Because it makes fewer comparisons than all-pairs testing, it has more power to detect those specific differences.
Holm-Šidák: a more powerful alternative to Bonferroni
Bonferroni divides the significance threshold by the number of comparisons. It is simple but conservative. Holm-Šidák applies the correction step by step, starting from the smallest p-value, and keeps the same protection against any false positive while rejecting more real effects.
Benjamini-Hochberg: controlling the false discovery rate
The methods above control the chance of making even one false positive. Benjamini-Hochberg instead controls the expected proportion of false positives among your significant results. It suits experiments with many comparisons, such as timecourses with many timepoints or screens, where some false discoveries are acceptable in exchange for power.
What to report
- the correction method used
- the number of comparisons it corrected for
- adjusted p-values, not the uncorrected ones
Join the private beta
All posts
Figures, 4 min read
SEM vs SD error bars: which should you show in your figures?
They look alike on a bar chart but answer different questions.
Standard deviation describes your data
The SD measures how spread out individual values are around the mean. It stays roughly the same whether you measure five mice or fifty. Show SD when you want readers to see the variability between samples.
Standard error describes your estimate of the mean
The SEM is the SD divided by the square root of n. It measures how precisely you have estimated the mean, so it shrinks as sample size grows. Four times as many samples gives an SEM half as large, even if the data are just as variable.
Common misreadings
- SEM bars are always smaller than SD bars, which can make variable data look tidy.
- Overlapping or non-overlapping SEM bars do not tell you whether a difference is significant.
- Error bars are meaningless if the legend doesn't say which kind they are and what n is.
A better habit
Show the individual data points on top of the bars. Readers can then judge spread and sample size directly, whichever error bar you choose. If your aim is to show the precision of a difference, a 95% confidence interval is often clearer than either.
Whatever you choose, state it in every figure legend, for example: "Data are mean ± SEM; n = 8 mice per group."
Join the private beta
All posts
Reporting, 4 min read
Why you shouldn't report post hoc power, and what to report instead
Calculating power after an experiment is tempting when a result isn't significant. It doesn't answer the question you are asking.
The problem
"Observed" or post hoc power uses the effect size you measured to compute how likely your study was to detect it. But that power is a direct function of the p-value you already have. A non-significant result will always come with low observed power, and a significant one with high power. It adds no new information, a point made clearly by Hoenig and Heisey in The American Statistician (2001).
Why it misleads
Low observed power is often used to argue that "the effect is probably real but the study was too small". Because the calculation is circular, it can't support that conclusion.
What to report instead
- Effect sizes with confidence intervals. A wide interval that spans both zero and a meaningful effect tells readers the study was inconclusive, and by how much.
- An a priori power analysis. Decide the sample size before the experiment, based on the smallest effect that would matter and a realistic estimate of variability, and report that plan.
This is why AnalyEase shows effect sizes and confidence intervals with every result, and never displays post hoc power.
Join the private beta