VVUQ · Module 6 of 10
Sources and Classification of Uncertainty
Not all uncertainty is the same. Some is the natural scatter of the world, which no amount of study removes; some is our own lack of knowledge, which more data can reduce. Telling them apart decides what to do next.
Readiness check
Learning objectives
- Classify an uncertainty component as aleatory or epistemic within a stated model and analysis scope.
- Distinguish model-form error from model-form uncertainty.
- Identify which mechanism — learning more, explaining more, or changing the system — could change a given component.
- Combine compatible standard uncertainties by root sum of squares under a stated zero-covariance assumption.
This module classifies uncertainty. Tick only what you can do closed-notes.
- Combine independent standard uncertainties as a root sum of squares.
- Recall that measured properties vary from sample to sample.
- Distinguish variation you can observe from knowledge you are missing.
- Recall a percentage reduction between two values.
- Recall that a model itself is an approximation.
The core idea
Aleatory uncertainty represents variability treated as random and irreducible by additional information within the current model and analysis scope. Epistemic uncertainty represents incomplete knowledge or representation that may, in principle, be reduced through additional evidence, improved modelling, or better characterization. Classifying each component tells you which kinds of effort could change it.
aleatory: variability treated as random in this modelepistemic: incomplete knowledge or representationutotal = √(ualeatory2 + uepistemic2)Uncertainty quantification begins by naming where the uncertainty comes from and what kind it is. Aleatory uncertainty is variability treated as random in the current model: the scatter in material strength between nominally identical parts, the turbulence in a flow, the roughness of a surface. It is usually characterised by a probability distribution. Epistemic uncertainty is incomplete knowledge or representation: a poorly known parameter, an untested boundary condition, or model-form uncertainty from approximations in the equations themselves.
The classification applies to a particular uncertainty component within a defined model and analysis. A physical source can contain both aleatory and epistemic components, and its decomposition may change when additional explanatory variables, evidence, or modelling detail are introduced. To classify a source, identify what varies, what is unknown, what the current model conditions on, what additional information could improve, and what intervention could change the physical variability.
“Irreducible” does not mean that the physical variability can never be altered. Process control, redesign, tighter tolerances, or conditioning on additional explanatory variables may change or explain the variability. It means that additional information alone does not remove the residual variability represented as random in the adopted model. Three mechanisms change an uncertainty, and they are worth keeping apart: learning more reduces uncertainty in parameters, distributions, inputs, or model adequacy; explaining more introduces conditioning variables or improved models that account for previously unexplained variation; changing the system uses process control, design changes, tighter tolerances, or operating restrictions to alter the physical variability itself.
Epistemic-dominant uncertainty often motivates additional evidence or model improvement. Aleatory-dominant variability often motivates characterization, robust design, margins, control, or process changes. The appropriate action depends on cost, sensitivity, and the decision.
The skills, taught in order
Five skills build the classification and its consequences.
6.1 Aleatory uncertainty
Aleatory uncertainty is variability treated as random and irreducible by additional information within the current model and analysis scope: the natural scatter in loads, material properties, geometry, and environment. It is characterised, usually by a probability distribution estimated from data. Additional observations do not reduce population variability, but they can reduce uncertainty in estimates of the population mean, variance, and other distribution parameters. Process control, tighter tolerances, redesign, or conditioning on an explanatory variable may change or explain the variability itself.
6.2 Epistemic uncertainty
Epistemic uncertainty is incomplete knowledge or representation that may, in principle, be reduced through additional evidence, improved modelling, or better characterization: an imprecisely known parameter, an assumed boundary condition, or a missing physical effect. It flags where investment in data or modelling may pay off. In practice a component may be reduced substantially without being removed entirely.
| Component | Origin | Reduced by more information? | Actions that may help |
|---|---|---|---|
| Aleatory | variability treated as random in this model | Not the variability itself; data can still sharpen its distribution parameters | characterise, robust design, margins, process control, conditioning variables |
| Epistemic (parameter) | poorly known input | Yes, in principle | measure it better |
| Epistemic (model form) | approximate equations | Often, but not always completely or practically | improve the model, gather validation evidence |
The two kinds of uncertainty component and the actions each opens up. Which dominates narrows the options; cost, sensitivity, and the decision determine the choice.
6.3 Model-form uncertainty
Two terms need separating. Model-form error is the actual inadequacy introduced by the mathematical model's representation of the physical system: a turbulence closure, a constitutive law, a neglected effect. Model-form uncertainty describes uncertainty about the magnitude or consequences of that inadequacy. Model-form uncertainty is generally epistemic, but it may not be completely or practically reducible. It is the hardest component to quantify, and validation evidence informs it, subject to the limits Modules 4 and 5 established: a discrepancy does not identify its own cause, uncertainty magnitudes are not known signed contributions, contributions can cancel, and validation evidence is scoped to the conditions tested.
6.4 Which mechanism could change it
Classification narrows the options rather than dictating one. An epistemic-dominant component often motivates additional evidence or model improvement. An aleatory-dominant component often motivates characterization, robust design, margins, control, or process changes. Ask which of the three mechanisms applies: learning more, explaining more with a richer model, or changing the system itself. The appropriate action depends on cost, sensitivity, and the decision the analysis supports.
6.5 Combining sources
In the simplified examples in this module, ua and ue are supplied as compatible standard-uncertainty contributions to the same output and are assumed independent. Under those assumptions, utotal = √(ua2 + ue2). The aleatory or epistemic label alone does not determine whether quadrature is valid: the uncertainty basis, output definition, units, dependence, and potential double counting must also be compatible, as Module 5 established.
Epistemic uncertainty is not always represented by a probability distribution or a standard deviation. When representations differ, aleatory and epistemic components may need to be propagated and reported separately. Probability distributions are often used to represent aleatory variability, but probability can also represent epistemic state of knowledge. Classification and mathematical representation are related modelling decisions, not automatic synonyms.
Keeping the parts separate within the total shows which mechanisms could change which portion.
6.6 Hybrid sources
Many engineering uncertainty sources are hybrid. Material strength may vary genuinely from specimen to specimen, while the distribution describing that variation is estimated from limited data. The specimen-to-specimen variability is treated as aleatory; uncertainty in the distribution parameters is epistemic. More tests sharpen knowledge of the distribution but do not make the specimens identical.
The same split appears elsewhere. Environmental loads vary genuinely from year to year, while a short historical record leaves the extreme-value behaviour imperfectly known. A measurement result contains repeatability effects that vary between readings alongside systematic corrections whose values are uncertain. In each case, asking which part is variability and which is imperfect knowledge points at different actions.
Engineering connection: deciding whether to run more tests, tighten a process, or add design margin turns on which component dominates, what each mechanism could change, and what the decision is worth.
Worked example 1: combining aleatory and epistemic uncertainty
A predicted quantity has an aleatory contribution ua = 0.04 (specimen-to-specimen material scatter) and an epistemic contribution ue = 0.03 (a poorly known boundary condition). Both are supplied as compatible standard uncertainties on the same output. Find the total.
- ProblemFind the total uncertainty from the sources in Figure 1.
- Given / findua = 0.04, ue = 0.03. Find utotal.
- AssumptionsThe two contributions are compatible standard uncertainties for the same output, in the same units, and are assumed independent, so every covariance term is taken as zero. The label alone does not license the combination; the basis does.
- Modelutotal = √(ua2 + ue2).
- Equationsutotal = √(0.042 + 0.032)
- Solveutotal = √(0.0016 + 0.0009) = √0.0025 = 0.05.
- CheckThe 3-4-5 combination gives exactly 0.05. The aleatory part (0.04) is the larger contributor, so most of the total would not be removed by additional information alone.
- ConclusionThe total is 0.05. The aleatory 0.04 would not be reduced by more information alone, though process control, tighter tolerances, or a richer model could change or explain that variability. Knowing the split is the point of the classification.
Worked example 2: how much can be reduced
For that result (utotal = 0.05, with ua = 0.04 aleatory and ue = 0.03 epistemic), consider the idealized limiting case in which the epistemic contribution is reduced to a negligible value while the aleatory contribution remains 0.04. Find the remaining total and the percentage reduction.
- ProblemFind the remaining uncertainty and reduction after removing the epistemic part, as in Figure 2.
- Given / findutotal = 0.05, ua = 0.04, ue = 0.03. Find the new total and the percentage reduction.
- AssumptionsAn idealized limiting case: the epistemic contribution is reduced to a negligible value and the aleatory contribution is unchanged. This is an upper bound on the benefit, not a description of what investigations usually achieve.
- ModelWith ue = 0, the remaining uncertainty is ua; reduction = (utotal − ua)/utotal.
- Equationsunew = ua = 0.04reduction = (0.05 − 0.04)/0.05
- Solveunew = 0.04. Reduction = (0.05 − 0.04)/0.05 = 0.20 = 20%.
- CheckAlthough the original epistemic contribution was 0.03, which is 75% of the aleatory 0.04, reducing it to a negligible value cuts the total by only 20%. A root sum of squares makes the effect of the smaller contributor on the total smaller than a linear comparison might suggest.
- ConclusionIn this limiting case the total falls from 0.05 to 0.04, a reduction of 20%. This is an upper-bound illustration of the possible benefit of reducing that epistemic contribution, not a claim that real investigations normally remove epistemic uncertainty completely.
Misconceptions and diagnostics
| Mistake | Symptom | Diagnostic question | Correction |
|---|---|---|---|
| Expecting data alone to shrink variability | Testing that never narrows the spread | "Would more data change the variability, or only my knowledge of its distribution?" | Characterise it. More data can sharpen the distribution; design, tolerances, or process control may change the variability itself. |
| Accepting epistemic gaps | A component that evidence could improve, left unaddressed | "Could evidence or a better model shrink this?" | Weigh investigation against cost, sensitivity, and the decision. |
| Ignoring model-form uncertainty | Only inputs counted | "Are the equations themselves approximate?" | Include model-form error from validation. |
| Combining incompatible quantities | A total that mixes bases | "Are these compatible standard uncertainties for the same output?" | Check basis, units, output definition, and dependence before combining; the label alone does not license it. |
Practice ladder
Classify each, then say what could change it: (a) scatter in yield strength, (b) an unknown friction coefficient, (c) a turbulence-model approximation.
Show answer
(a) aleatory in most analyses, though the distribution parameters are estimated from a finite sample and that part is epistemic; tighter process control or a supplier change could alter the variability itself. (b) epistemic (parameter): measurement can reduce it. (c) epistemic (model form): a better closure or validation evidence can reduce it, though perhaps not completely.
Combine ua = 0.06 and ue = 0.08, given as compatible standard uncertainties for the same output and assumed independent.
Show answer
utotal = √(0.062 + 0.082) = √0.01 = 0.10.
For that case, how much would the total fall in the limiting case where the epistemic 0.08 becomes negligible?
Show answer
Remaining = ua = 0.06. Reduction = (0.10 − 0.06)/0.10 = 40%. The epistemic contribution was the dominant one here, so reducing it helps more than in Worked Example 2. This is again an upper bound on the benefit.
A prediction's uncertainty is dominated by a poorly known material parameter. Work through: what is physically variable; what is imperfectly known; whether the source contains both components; what more data could improve; what a richer model could explain; what design or process control could change; and how the uncertainty is represented in this analysis. Then argue for the next action.
What good work looks like
The dominant component is epistemic, so measuring the parameter more precisely is the action most likely to shrink the total. A good answer also notes that the parameter may itself vary between batches, which is aleatory and would not vanish with better measurement; that the representation chosen for the parameter affects how it propagates; and that the case for testing depends on its cost against the sensitivity of the decision.
Working with AI, and proving it yourself
Use AI as an examiner, not a solver
Portfolio task
For a real prediction, list the uncertainty sources; for each, say what varies, what is imperfectly known, and whether it is hybrid; state how each is represented; combine those on a compatible basis; and identify which component to target and by which mechanism.
Retrieval and spaced review
Closed notes. Answer out loud, then reveal.
1. What is aleatory uncertainty?
Variability treated as random and irreducible by additional information within the current model and analysis scope. Data can still sharpen its distribution; design or process changes can alter the variability itself.
2. What is epistemic uncertainty?
Incomplete knowledge or representation, in parameters or model form, that may in principle be reduced by evidence or better modelling, though not always completely.
3. What is model-form uncertainty?
Uncertainty about the magnitude or consequences of the model's representational inadequacy. The inadequacy itself is model-form error.
4. When may the two be combined as a root sum of squares?
When they are compatible standard uncertainties for the same output and are treated as independent. The label alone does not license it.
5. Why does the classification matter?
It narrows which mechanisms could change the uncertainty: learning more, explaining more with a richer model, or changing the system.