Balanced crossed factorial experiment
Two Way ANOVA Calculator
Partition variation in a replicated, balanced two-factor design. Enter one line per treatment combination, view main effects and interaction separately, and audit the complete sums-of-squares identity instead of accepting a single unexplained p-value.
Build the factorial cell matrix
Factorial variance table
The default example shows evidence for Lab and Material main effects at α = 0.05, but not for their interaction.
| Source | SS | df | MS | F | p-value |
|---|---|---|---|---|---|
| Lab | 5.0139 | 1 | 5.0139 | 100.2778 | < 0.0001 |
| Material | 2.1811 | 2 | 1.0906 | 21.8111 | 0.0001 |
| Interaction | 0.1344 | 2 | 0.0672 | 1.3444 | 0.2969 |
| Error | 0.6000 | 12 | 0.0500 | — | — |
| Corrected total | 7.9294 | 17 | — | — | — |
Mean map and partition check
Grand mean: 3.0056 Lab means: Lab 1 = 3.5333; Lab 2 = 2.4778 Material means: Material 1 = 3.4500; Material 2 = 2.6000; Material 3 = 2.9667 Replicates per cell: 3; observations: 18 SS check: 5.0139 + 2.1811 + 0.1344 + 0.6000 = 7.9294
What this two-way ANOVA calculates
A two-way analysis of variance examines a quantitative response across every combination of two categorical factors. This calculator implements a balanced, crossed, fixed-effects design with replication. “Crossed” means each level of Factor A appears with each level of Factor B. “Balanced” means every combination has the same number of observations. “With replication” means every cell has at least two observations, which supplies a within-cell error estimate and permits a separate interaction test.
The model expresses each observation as a grand mean, a Factor A effect, a Factor B effect, an A-by-B interaction effect, and random error. The calculator partitions corrected total sum of squares into those four explainable pieces. It then divides each effect mean square by the error mean square to form an F statistic. Each right-tail F probability is reported as a p-value.
This scope is deliberately narrower than a general statistical package. It does not fit unbalanced Type I, II, or III sums of squares; random effects; nested factors; repeated measures; covariates; missing cells; or more than two factors. A clear limitation prevents a balanced classroom formula from being silently applied to a different design.
Read interaction before main effects
Interaction asks whether effects change
An interaction means the effect of one factor depends on the level of the other. For example, if a coating performs better on Material 1 in Lab 1 but worse on the same material in Lab 2, a single average lab difference can hide the pattern. Inspect cell means and an interaction plot when the interaction is important.
Main effects average over the other factor
The Factor A test compares A-level marginal means averaged across B. The Factor B test does the corresponding comparison averaged across A. Those averages are meaningful only in the context of the observed interaction. A main effect can be statistically detectable even when particular cells move in different directions.
No interaction evidence is not proof of parallelism
A large interaction p-value says the experiment did not detect interaction under its model; it does not demonstrate that every cell contrast is exactly equal. Small replication can leave interaction estimates imprecise. Use confidence intervals and design-relevant minimum effects in serious planning.
The variance partition in four layers
Factor A
Each A marginal mean is compared with the grand mean. Squared deviations are multiplied by the number of B levels and replicates.
Factor B
Each B marginal mean is compared with the grand mean. Squared deviations are multiplied by the number of A levels and replicates.
A × B interaction
Each cell mean is compared with what additive A and B effects would predict. The squared departures are multiplied by replicates.
Within-cell error
Every observation is compared with its own cell mean. This residual variation becomes the denominator for all three F tests.
The identity SS(total) = SS(A) + SS(B) + SS(AB) + SSE is printed with the result. Small last-digit differences can appear from display rounding, while the script retains full precision internally. The degrees of freedom also reconcile: N − 1 equals (a − 1) + (b − 1) + (a − 1)(b − 1) + ab(r − 1).
Worked example: coating results across labs and materials
The default data reproduce a NIST example with two laboratories, three materials, and three replicate coating measurements per lab-material cell. The grand mean is about 3.0056. Lab 1’s marginal mean is about 3.5333, while Lab 2’s is about 2.4778. Material marginal means are approximately 3.45, 2.60, and 2.9667.
The variance partition assigns about 5.0139 units of squared variation to Lab, 2.1811 to Material, 0.1344 to interaction, and 0.6000 to within-cell error. With error mean square 0.05, the Lab and Material F ratios are large. The interaction F ratio is only about 1.34. At a conventional 0.05 level, the example detects both main effects but does not detect interaction.
The output should be read as an analysis of this designed example, not as proof that every lab or material in a wider population behaves identically. Factor levels are treated as the fixed levels of interest. Generalizing to randomly sampled labs or materials calls for a random- or mixed-effects model and a different variance interpretation.
Input format and data audit
Write one nonblank line for each unique combination using three fields separated by vertical bars: Factor A level, Factor B level, and comma-separated response values. Labels may contain spaces but cannot be blank. Every response must be finite. Repeat counts must match across all cells, and the calculator requires at least two levels per factor and at least two replicates per cell.
Each combination must appear exactly once. Duplicate lines would create ambiguous cells, so they are rejected. A complete crossed matrix must contain a × b lines. If the labels reveal two A levels and three B levels, exactly six distinct cells must exist. These safeguards catch accidental omissions before the sums of squares are built.
Do not paste identifiers, notes, inequality symbols, units, or missing-value codes into the response list. Convert measurement units consistently and decide how exclusions are handled before calculation. Keep a protected copy of raw data, because this summary interface cannot trace an edited number back to its source record.
Assumptions and design limits
Classical fixed-effects ANOVA assumes independent errors, approximately normal errors within treatment combinations, and a common error variance. Balanced designs offer some robustness, but serious outliers, strong skewness with low replication, or markedly different cell variances can distort inference. Plot residuals against fitted values, examine normal-probability behavior, and investigate data provenance rather than relying on p-values alone.
Random assignment supports causal interpretation; merely observing preexisting groups generally does not. Blocking, repeated observations on the same unit, batches, classrooms, sites, and households can create dependence. Entering them as if independent inflates effective information. A statistician or suitable model should reflect the actual assignment and sampling structure.
If each cell contains only one observation, interaction and residual error cannot be separated by this replicated design. If cells have unequal counts, use statistical software that states its sums-of-squares convention and handles estimability. Never pad cells, duplicate observations, or discard valid data merely to satisfy a balanced calculator.
How to report the analysis
Name the response and both factors, identify the levels, state that the design was crossed and balanced, give replicates per cell and total N, and describe assignment and independence. Report the interaction first, then the factor tests with F degrees of freedom and exact p-values. Include cell means or an interaction plot and relevant confidence intervals. A compact example is: “A balanced 2 × 3 fixed-effects ANOVA with three replicates per cell found Lab and Material main effects, with no detected Lab × Material interaction.” Follow that sentence with the actual F and p values.
Avoid reducing results to three binary significance labels. State units, effect patterns, uncertainty, multiplicity decisions, diagnostic findings, exclusions, and whether the analysis was prespecified. When a factor has more than two levels and its omnibus test is important, planned contrasts or multiplicity-adjusted follow-up comparisons are usually needed to learn which levels differ. The omnibus F test alone does not identify pairs.
Frequently asked questions
Can I use unequal cell sizes?
No. This calculator intentionally implements the balanced replicated formulas. Use a linear-model package for unbalanced data and report the sums-of-squares convention.
What if I have only one value in each cell?
Without replication, the ordinary interaction and residual error cannot both be estimated separately. A no-replication analysis has different assumptions and is outside this calculator.
Does a significant interaction cancel main effects?
It does not erase the arithmetic, but it changes the question. Marginal main effects average across the other factor and may conceal level-specific directions, so interpret cell contrasts first.
Are the factor levels random or fixed?
The calculator treats the entered levels as fixed levels of interest. Inference about a broader population of randomly sampled levels requires a random- or mixed-effects model.
Can ANOVA establish causation?
Only a suitable design, such as credible random assignment with proper execution, supports causal claims. The calculation alone cannot remove confounding or selection bias.
What should follow a significant factor with three levels?
Use prespecified contrasts or multiplicity-aware comparisons that answer the scientific question. Do not run many unadjusted pairwise tests and report only favorable ones.
Related calculator
Matching the tool to the design is more important than forcing every mean comparison into an ANOVA table.
When the design has only one categorical factor, use the simpler one-way ANOVA statistical analysis instead of forcing a second factor.
References
The balanced fixed-effects formulas, degrees of freedom, and default coating example follow the U.S. National Institute of Standards and Technology, The two-way ANOVA, and its companion models and calculations. Those pages define the factorial model and the partition SS(total) = SS(A) + SS(B) + SS(AB) + SSE.