VVUQ · Module 4 of 10
Validation Experiments and the Validation Hierarchy
Validation compares a simulation to reality, but only over the conditions actually tested. A validation hierarchy builds trust from simple pieces up to the full system, and predictions are safe only inside the validated range.
Readiness check
Learning objectives
- Interpret what comparing the validation discrepancy |E| with the stated experimental uncertainty can and cannot establish.
- Identify how a validation hierarchy builds evidence from unit problems to the full system.
- Determine whether an application point is interpolative or extrapolative relative to the validation domain.
- Distinguish a validation experiment from an ordinary test.
This module compares models to experiments. Tick only what you can do closed-notes.
- Recall the comparison error E = S − D.
- Recall that a measurement has an uncertainty.
- Distinguish interpolation from extrapolation.
- Recall that validation evidence applies only at the conditions tested.
- Recall a dimensionless group such as the Reynolds number.
The core idea
Validation compares a verified simulation to experimental data designed for the purpose. A validation hierarchy builds evidence from unit problems up to the full system. Validation produces evidence for specified quantities under specified tested conditions; it does not certify a model.
compare S to D at tested conditionshierarchy: unit → benchmark → subsystem → systempredict inside the validated range, not beyondOnce a solution is verified, validation asks whether the model matches reality. This is done with validation experiments, tests designed specifically to measure the quantities a simulation predicts, with well-characterised conditions and quantified measurement uncertainty, not repurposed qualification tests. Because a full engineering system is too complex to validate in one step, validation is built as a hierarchy: unit-level problems isolate single physics, benchmark cases combine a few, subsystem tests add coupling, and only at the top does the complete system appear. Evidence accumulates up the pyramid, so confidence in a system prediction rests on validated pieces beneath it. The comparison discrepancy E = S − D is the starting point, but interpreting it requires the uncertainty in the comparison. Experimental uncertainty is only one contributor. Numerical uncertainty from Module 3, input uncertainty, and any relevant dependence or correlation must also be considered, and Module 5 develops the quantitative combination. Validation produces evidence for specified quantities under specified tested conditions. A prediction between tested conditions is interpolative in that coordinate, though applicability still has to be assessed; a prediction beyond them is extrapolative and is not covered by the available evidence.
The skills, taught in order
Five skills build validation experiments, the hierarchy, and the limits of a validation.
4.1 Validation experiments
A validation experiment is designed to test a model, not to qualify a product: it measures the predicted quantities under controlled, well-characterised conditions with quantified uncertainty. Generate the simulation prediction before examining or using the experimental response measurements that will serve as validation evidence. Measured as-run inputs, geometry, boundary conditions, and operating conditions may still be needed to reproduce the actual experiment. The distinction is that validation response data must not be used to calibrate or tune the model being evaluated.
4.2 Comparing simulation to data
Work in four steps. First, calculate the comparison discrepancy E = S − D. Second, identify the uncertainty sources needed to interpret it. Third, recognise that experimental uncertainty alone is incomplete: here uD is the experimental standard uncertainty, and comparing |E| with it accounts for neither numerical nor input uncertainty, nor any dependence between contributors, and fixes no coverage factor or probability model. Fourth, carry E and its uncertainty sources into Module 5. Module 5 combines the numerical, input, and experimental contributions after defining them on a compatible standard-uncertainty basis. The root-sum-square expression uval = √(unum² + uinput² + uD²) applies when unum, uinput, and uD are standard uncertainties and the relevant error contributions are treated as independent. If dependence is important, covariance terms are required. Note that the GCI reported in Module 3 is a spatial-discretization uncertainty estimate constructed with a safety factor. It is not automatically identical to the standard uncertainty unum used in a validation-uncertainty combination; that relationship must be defined explicitly, which Module 5 does before using the combination. Comparing |E| with uD alone supports no verdict about model-form discrepancy.
| Level | What it tests | Physics |
|---|---|---|
| Unit problem | a single effect | isolated |
| Benchmark | a few effects | partly coupled |
| Subsystem | a component | coupled |
| System | the full application | fully coupled |
The validation hierarchy. Simpler levels isolate physics and are easier to measure precisely; the system level is the goal but the hardest to test.
4.3 The validation hierarchy
Because a full system cannot be validated in one experiment, evidence is built up a hierarchy from unit problems to the complete application. Each level validates the physics it adds, so a credible system prediction stands on a foundation of validated pieces.
4.4 The validation domain
Validation evidence covers only the conditions actually tested, which bound the validation domain. Predictions at conditions between tested points are interpolation, supported by the validation; predictions outside are extrapolation, which the validation does not cover.
4.5 Interpolation versus extrapolation
Evidence obtained at tested conditions does not automatically transfer to an untested condition, even one that is interpolative: applicability depends on the quantity of interest, other operating conditions, model assumptions, and whether the governing physical regime remains comparable. Using the model outside the tested range requires additional justification, because the physics may change and no data constrain the model there. Recognising which side of the boundary a prediction falls on is essential to credibility.
Engineering connection: a turbulence model validated at one Reynolds number range cannot be assumed valid at another without new evidence, a distinction that governs every simulation-based decision.
Worked example 1: what the discrepancy alone tells you
A simulation predicts S = 1.05 and a validation experiment measures D = 1.00 with an experimental standard uncertainty uD = 0.03. Calculate the comparison discrepancy and state what this comparison does and does not establish.
- ProblemCalculate the comparison discrepancy in Figure 1 and state what it establishes.
- Given / findS = 1.05, D = 1.00, and uD = 0.03 as an experimental standard uncertainty. Find E and its magnitude relative to uD.
- AssumptionsThe solution is verified. Numerical uncertainty and input uncertainty are not quantified here, so they are unknown rather than zero. No coverage factor, probability model, or bounding interpretation has been defined.
- ModelE = S − D. The ratio |E|/uD expresses the discrepancy magnitude in units of the stated experimental standard uncertainty. It is a descriptive ratio, not a decision rule.
- EquationsE = S − D = 0.05|E| / uD = 0.05 / 0.03 ≈ 1.67
- SolveE = 1.05 − 1.00 = 0.05, so |E|/uD = 0.05/0.03 ≈ 1.67. The discrepancy magnitude is about 1.67 times the stated experimental standard uncertainty.
- CheckAsk what is missing before drawing any conclusion. Numerical uncertainty from solution verification and uncertainty in the simulation inputs are absent, correlation between contributors has not been considered, and no coverage factor or probability model has been stated. Each of those could change the interpretation in either direction.
- ConclusionThe observed simulation-experiment discrepancy is E = 0.05, which is approximately 1.67 times the stated experimental standard uncertainty. Numerical uncertainty, input uncertainty, correlation effects, and an interpretation or coverage criterion have not yet been included. Therefore, this comparison alone does not establish whether a model-form discrepancy has been resolved or how large such a discrepancy might be.
Worked example 2: inside or outside the validation domain
Validation comparisons were performed at Reynolds numbers 1000 (comparison discrepancy 2%) and 5000 (3%). Classify predictions at Re = 3000 and at Re = 10000 relative to the available validation points.
- ProblemClassify the two predictions in Figure 2 relative to the available validation points.
- Given / findValidation evidence is available at Re = 1000 (2%) and Re = 5000 (3%). Classify Re = 3000 and Re = 10000.
- AssumptionsTwo validation points only. Whether the governing physical regime stays comparable between them has not been demonstrated, and no data exist beyond them.
- ModelClassify each condition relative to the tested points: between them is interpolative in that coordinate, beyond them is extrapolative. Classification is geometric; applicability is a separate judgement.
- Equationsvalidation points at Re = 1000 and Re = 5000
- SolveRe = 3000 lies between the two points, so it is interpolative with respect to Reynolds number. Applicability of the evidence there also depends on the quantity of interest, other operating conditions, model assumptions, and whether the governing physical regime remains comparable. Re = 10000 lies beyond both points and is extrapolative with respect to Reynolds number relative to the available validation points.
- CheckTwo comparisons do not establish a continuous validated range between them. The 2% and 3% discrepancies belong to their own conditions and cannot be transferred to Re = 3000, and the flow regime could shift well before Re = 10000.
- ConclusionRe = 3000 is interpolative and may be supportable subject to an applicability assessment. Re = 10000 is extrapolative and is not covered by the available evidence. Classification is the first question, not the last.
Misconceptions and diagnostics
| Mistake | Symptom | Diagnostic question | Correction |
|---|---|---|---|
| Ignoring data uncertainty | Small gap called a discrepancy | "Which uncertainties have I included?" | Express |E| relative to the stated uncertainties, and name the contributors still missing. |
| Extrapolating a validation | Trusting the model out of range | "Is this inside the span of validation points?" | Validation evidence applies at the conditions tested; transfer elsewhere needs an applicability assessment. |
| Tuning to the data | Model adjusted after seeing D | "Was S predicted blind?" | Predict before comparing to avoid calibration-as-validation. |
| Validating only the system | No supporting unit evidence | "Are the lower levels validated?" | Build the hierarchy from unit problems up. |
Practice ladder
S = 98, D = 100, with an experimental standard uncertainty uD = 5. State what this comparison establishes.
Show answer
E = 98 − 100 = −2, so |E| = 2 and |E|/uD = 0.4. The discrepancy magnitude is smaller than the stated experimental uncertainty. This limited comparison does not resolve a discrepancy relative to that uncertainty and does not yet include the other contributors to validation uncertainty.
Order these validation levels from base to peak: subsystem, unit, system, benchmark.
Show answer
Unit, benchmark, subsystem, system, from isolated physics up to the full application.
Validation comparisons were performed at loads of 10 kN and 50 kN. Classify predictions at 30 kN and 70 kN relative to those points.
Show answer
30 kN lies between the two validation points, so it is interpolative with respect to load, subject to an applicability assessment. 70 kN lies beyond both points and is extrapolative relative to the available validation evidence.
You must trust a system-level simulation for a new operating point. Describe the validation evidence you would want beneath it.
What good work looks like
Validated unit and benchmark cases for each key physics, subsystem tests for the couplings, and system data bracketing the operating point, so the prediction is interpolation within a hierarchy of validated evidence, not an unsupported extrapolation.
Working with AI, and proving it yourself
Use AI as an examiner, not a solver
Portfolio task
For a real validation, plot simulation against data with the measurement uncertainty, state the validation domain, and classify an intended prediction as interpolation or extrapolation.
Retrieval and spaced review
Closed notes. Answer out loud, then reveal.
1. What is a validation experiment?
A test designed to measure the quantities a model predicts, with quantified uncertainty.
2. What does comparing |E| with uD alone establish?
Only the discrepancy magnitude relative to the stated experimental standard uncertainty. It excludes numerical and input uncertainty, ignores dependence between contributors, and sets no coverage or probability interpretation, so it supports no verdict on model-form discrepancy.
3. What is the validation hierarchy?
Evidence built from unit problems up through subsystems to the full system.
4. What bounds the validation domain?
The conditions actually tested. Two comparisons do not by themselves establish a continuous validated range between them.
5. Interpolation versus extrapolation?
A condition between tested points is interpolative in that coordinate; beyond them it is extrapolative. Classification is geometric, and applicability is a separate judgement in both cases.