Bayes Theorem Calculator
Update a prior probability after positive or negative evidence using sensitivity and specificity. See the same calculation as probabilities, odds, likelihood ratios, and natural frequencies so the base-rate effect is visible rather than buried in one posterior percentage.
In a cohort of 10,000, the model expects 90 true positives and 495 false positives.
Describe the hypothesis and evidence
Prior and evidence performance
Prior sensitivity analysis
These two values reuse the same sensitivity and specificity to show how base rate changes interpretation.
Posterior evidence map
0.0101
18.00
0.1818
Bayes updates probabilities under the entered model. It does not validate the prior, prove sensitivity and specificity apply to this population, or choose an action threshold.
What Bayes’ theorem answers
Bayes’ theorem updates the probability of a hypothesis after observing evidence. The prior probability describes belief or prevalence before the evidence. Sensitivity is the chance of positive evidence when the hypothesis is true. Specificity is the chance of negative evidence when the hypothesis is false. The posterior is the revised probability.
The question “How often is the evidence positive when H is true?” is not the same as “How often is H true when the evidence is positive?” Bayes reverses that conditioning by including the base rate. Confusing the two is the inverse-probability fallacy.
Positive-evidence equation
False-positive rate P(E|not H) = 1 − specificity
With a 1 percent prior, 90 percent sensitivity, and 95 percent specificity, the numerator is 0.90 × 0.01 = 0.009. The false-positive component is 0.05 × 0.99 = 0.0495. Dividing 0.009 by 0.0585 produces 0.153846, or 15.38 percent.
Negative-evidence equation
Under the default, 10 percent of true cases produce negative evidence. That contributes 0.001 of the population, while true negatives contribute 0.9405. The posterior after a negative is about 0.106 percent. A negative result lowers the prior but does not make the probability zero.
Natural frequencies make the base rate visible
Imagine 10,000 comparable cases. A 1 percent prior means 100 have H and 9,900 do not. Sensitivity finds 90 of the 100, leaving 10 false negatives. A 5 percent false-positive rate marks 495 of the 9,900 non-H cases positive and leaves 9,405 true negatives.
Among 585 expected positives, only 90 are true positives. Thus 90 ÷ 585 equals the 15.38 percent posterior. Counts may be fractional for another cohort size because they are expected frequencies, not a promise that a real finite sample produces exact integers.
Odds and likelihood ratios
Probability p converts to odds p ÷ (1 − p). Positive likelihood ratio is sensitivity divided by false-positive rate. Default LR+ is 0.90 ÷ 0.05 = 18. Prior odds of 0.0101 multiplied by 18 become posterior odds of 0.1818, which converts back to 15.38 percent probability.
Negative likelihood ratio is false-negative rate divided by specificity: 0.10 ÷ 0.95, or about 0.1053. Odds form is convenient for sequential evidence, but multiplying likelihood ratios assumes the ratios are appropriate for the case and conditional dependence is handled correctly.
Why prevalence changes positive predictive value
The same test performance yields very different posteriors at different priors. The default sensitivity strip reports about 8.29 percent after a positive at 0.5 percent prior, 15.38 percent at 1 percent, and 48.65 percent at 5 percent. A rare hypothesis creates many opportunities for false positives among the much larger non-H group.
Sensitivity and specificity are not predictive values. Positive predictive value depends on prior probability in the target population. Reporting “90 percent accurate” without defining population, threshold, reference standard, and metric is inadequate.
Where a prior comes from
A prior may come from a representative prevalence study, historical process frequency, calibrated forecasting model, previous experiment, or explicitly stated belief distribution. It should match the population, time, threshold, and hypothesis. Convenience data can produce a biased prior.
For an individual medical question, population prevalence may need adjustment for symptoms, age, exposure, history, and selection into testing. That adjustment is clinical work, not a number this general calculator can perform. Do not substitute the default example for professional assessment.
Sensitivity and specificity are conditional too
Published test performance may vary by disease stage, specimen handling, threshold, equipment, reader, reference standard, and population spectrum. Sensitivity and specificity estimated in an enriched case-control study may not transport unchanged to screening practice. Confidence intervals matter.
Use values from a valid study that resembles the intended use. If uncertainty is large, run low and high performance scenarios rather than copying point estimates with false precision. Bayes computes correctly only for the inputs and model supplied.
Sequential evidence and dependence
The first posterior can become the next prior, but evidence sources may overlap. Two tests measuring the same underlying feature are not automatically independent. Multiplying their likelihood ratios as though independent can exaggerate certainty. Use joint performance data or a model of dependence when available.
Repeated measurements can also share systematic error. Ten copies of one biased sensor are not ten independent confirmations. Document what each evidence item contributes and whether its performance was measured after conditioning on earlier evidence.
Medical testing caution
This tool can illustrate diagnostic-test arithmetic but cannot diagnose, rule out disease, select a test, or recommend treatment. A result may require confirmatory testing, clinical context, urgency assessment, and discussion of harms from false positives and false negatives. Seek qualified care for medical decisions.
NCBI’s diagnostic-testing resources emphasize that posttest probability depends on pretest probability plus test sensitivity and false-positive performance. They also describe limitations when performance differs across patient subgroups or sequential tests are dependent.
Fraud, spam, and quality-control examples
Outside medicine, H might mean a transaction is fraudulent, a message is spam, or a manufactured unit is defective. Sensitivity becomes the detection rate; specificity becomes the correct-clear rate. The cost of mistakes may influence the decision threshold even when posterior probability is unchanged.
A fraud model can be mathematically well calibrated yet create unfair or unlawful outcomes if inputs reflect biased data or actions ignore due process. Bayes describes probability updating, not ethics, legality, fairness, or the optimal intervention.
Posterior probability is not a decision rule
Action depends on consequences, costs, benefits, alternatives, delay, reversibility, and risk tolerance. A 15 percent posterior may be high enough for a cheap follow-up test but far below the threshold for an irreversible action. Thresholds should be set before observing a convenient result when possible.
Do not transform probability into certainty through language. Report assumptions, posterior, evidence direction, and sensitivity range. If a downstream calculator uses the posterior, preserve more precision internally while displaying sensible rounding.
Calibration versus discrimination
Discrimination separates H from non-H cases; calibration checks whether events occur at predicted rates. Sensitivity and specificity at one threshold do not establish overall calibration. A score can rank cases well while its stated probabilities are too high or low.
Validate predictions on new data, inspect subgroups, and update when prevalence or process changes. A posterior built from stale performance may look exact while being poorly calibrated. Natural-frequency audits help stakeholders see what the model implies.
Using the calculator responsibly
Define H and E before entering values. Confirm that sensitivity and specificity use the same positive threshold and reference truth. Record the source and uncertainty. Run alternative priors and performance bounds. Compare expected true/false counts and decide whether the evidence meaningfully changes an action.
Do not carry medical-test labels or assumptions into game probability, and do not treat independent-trial formulas as evidence that an outcome is “due.”
Audit the denominator, not only the numerator
Many incorrect Bayes calculations multiply prior by sensitivity and stop. That numerator counts true-positive probability but does not answer how many total positives exist. The denominator must add true positives and false positives. Write both terms with labels before substituting numbers, then confirm they sum to the overall positive-evidence rate shown in the result.
Also check complements deliberately: non-H is one minus prior, false-positive rate is one minus specificity, and false-negative rate is one minus sensitivity. A spreadsheet that subtracts sensitivity where it should subtract specificity can produce a plausible-looking but wrong posterior. The natural-frequency grid is a practical cross-check because its four cells must sum to the cohort and its positive cells must reproduce the posterior.
Frequently asked questions
Is sensitivity the probability that H is true after a positive?
No. Sensitivity is P(positive|H). The posterior P(H|positive) also depends on prior probability and false positives.
Can a highly sensitive test have a low positive posterior?
Yes. If H is rare or specificity is insufficient, false positives can outnumber true positives.
Why are expected cohort counts fractional sometimes?
They are probability-based expectations. A real finite cohort has integer outcomes that vary around them.
Can I update with two tests?
Only with appropriate conditional likelihoods or justified independence. Blindly multiplying correlated evidence overstates certainty.
Does a negative result make the posterior zero?
Only under special assumptions such as perfect sensitivity. Normally it reduces probability by the negative likelihood ratio.
Is the posterior an instruction to act?
No. Decisions require consequences, alternatives, thresholds, uncertainty, and domain expertise beyond probability.
References
NCBI Bookshelf — The Use of Diagnostic Tests: A Probabilistic Approach
NCBI Bookshelf — The Diagnostic Process
NCBI Bookshelf — Sensitivity, Specificity, and Predictive Value