VVUQ · Module 10 of 10

Credibility Assessment and Decision-Making

All the verification, validation, and uncertainty evidence exists to serve one thing: a decision. This closing module scales the required rigor to what is at stake and turns the evidence into a documented judgement of whether it meets, partly meets, or does not meet the credibility goals for a stated use.

01

Readiness check

Learning objectives

  • Evaluate model risk from model influence and decision consequence as levels, without computing a score.
  • Distinguish evidence relevance from evidence adequacy.
  • Calculate the normalized margin RM = (L − Sapp)/upred and interpret it against a pre-established, decision-specific criterion on a declared uncertainty basis.
  • Document a credibility record that states residual limitations and the decision owner.

This closing module turns evidence into decisions. Tick only what you can do closed-notes.

  • Recall the validation uncertainty and prediction uncertainty.
  • Judge two separate factors and read them together, without combining them into one number.
  • Compute a margin as a difference divided by an uncertainty.
  • Recall that model risk here combines model influence and decision consequence, not the probability that the model is wrong.
  • Recall that a model has an intended use.
0 or 1 weak itemsContinue with this module.
2 weak itemsRevisit prediction uncertainty in Module 9.
3 or more weak itemsRevisit the framework in Module 1.
02

The core idea

Credibility is VVUQ evidence that is both relevant and adequate for a specific decision. The ASME V&V 40 approach scales the required rigor to the model's influence and the decision's consequence. Adequacy is judged against goals set for that context of use, and a quantitative check compares the prediction, with its uncertainty, against the limit and a decision-specific margin criterion.

model risk rises with influence AND with consequencehigher risk ⇒ more VVUQ evidenceRM = (L − Sapp) / upred, criterion RM ≥ 2 (decision-specific)

The point of verification, validation, and uncertainty quantification is not to produce numbers but to support a decision, and the amount of evidence required depends on what the decision is worth. The ASME V&V 40 credibility framework makes this risk-informed. ASME V&V 40-2018 was developed for medical-device applications; ASME considers its risk-informed framework transferable to other disciplines, but the standard is not a universal normative requirement for every mechanical-engineering analysis. This course uses its structure as a credibility-assessment example and relies on engineering judgement and applicable domain-specific requirements. V&V 40 is a framework for assessing the relevance and adequacy of completed V&V activities; it is not a quantitative formula, a universal scoring method, or a step-by-step prescription of the evidence required for every application. Model risk combines two factors: the model's influence on the decision (how much the outcome relies on the simulation rather than on tests or other evidence) and the consequence of the decision being wrong (how severe a failure would be). A high-influence, high-consequence use demands extensive VVUQ evidence, code and solution verification, validation on relevant data, and quantified uncertainty; a low-risk use can be credible with far less. Both factors are judged as levels against the stated context of use, and risk rises as either one rises. Neither the standard nor the FDA guidance that applies it gives a formula, a numeric scale, or a threshold that separates the bands, so any scoring scheme you meet is someone's local convention and should be labelled as one. The framework is set out in ASME V&V 40-2018, Assessing Credibility of Computational Modeling through Verification and Validation: Application to Medical Devices, and applied in the FDA's 2023 guidance Assessing the Credibility of Computational Modeling and Simulation in Medical Device Submissions, which describes model risk as rising with model influence and with decision consequence rather than as a computed number. Adequacy for intended use is then the judgement that the evidence matches the risk. Evidence is judged on two questions, not one: relevance asks whether an activity challenges the assumptions, mechanisms, quantities, and conditions that matter to the context of use, and adequacy asks whether that relevant evidence is rigorous and complete enough to meet the goal set for it. Evidence can be technically strong but irrelevant, or highly relevant but inadequate, and a critical gap is not offset by strong evidence elsewhere. A separate quantitative check compares the prediction against the limit: the normalized margin RM = (L − Sapp)/upred is the gap in units of the combined prediction uncertainty from Module 9, and a decision-specific criterion such as RM ≥ 2, set before the comparison, records how much margin the consequence of error demands. Report whether upred is a standard or an expanded uncertainty, because the criterion must be read on the same basis. Documenting the evidence, its relevance and adequacy, the uncertainty basis, the margin, the residual limitations, and who owns the decision is what makes a simulation-based decision defensible to a reviewer or regulator.

The skill works when: you scale the evidence to the risk and confirm the margin against the uncertainty.
The skill breaks down when: the same rigor is applied regardless of stakes, or a margin ignores the prediction uncertainty.
The concept. Model risk grows with both the model's influence on the decision and the consequence of being wrong. The required VVUQ evidence scales with that risk.
03

The skills, taught in order

Five skills turn the accumulated evidence into a risk-informed, documented decision.

10.1 Context of use and intended use

Credibility is not a property of a model in the abstract but of a model for a stated use. The question of interest is the specific decision question the analysis must answer; the quantity of interest is the physical response used to answer it. The context of use links these to the system configuration, application conditions, model output, the model's role alongside other evidence, the intended user, and the decision consequence. It is not the model's broad general purpose: the same simulation may be credible for a rough concept study and inadequate for a certification, so the context of use must be fixed before credibility is judged.

10.2 Model risk

Model risk combines the model's influence on the decision and the consequence of an incorrect decision. Influence is how strongly the decision relies on the computational result rather than on tests or other evidence; it is not the sensitivity of the output to an input. Consequence is the severity of a wrong decision, not the probability that it is wrong. High influence with high consequence is high risk; either being low lowers the risk. This risk sets how much VVUQ evidence is required.

Model riskEvidence expected
Lowbasic verification, limited validation
Moderatesolution verification, relevant validation, UQ
Highfull code and solution verification, extensive validation, quantified UQ

The V&V 40 principle: the rigor of the VVUQ evidence is scaled to the model risk, not applied uniformly. The levels are judged, not computed; influence and consequence are not multiplied or averaged into a number.

10.3 Credibility goals, relevance, and adequacy

For each credibility factor relevant to the context of use, establish a credibility goal describing the expected rigor or strength of evidence, and set it from the model risk before the evidence is judged, not from what the results turn out to be. The collection of factor-specific goals defines the credibility burden for that use; a goal is not a prediction that the evidence will pass, and it should not be moved after weak evidence is found.

Then judge each item of evidence on two separate questions. Relevance asks whether the evidence challenges the assumptions, physical mechanisms, quantities, and conditions that matter to the context of use. Adequacy asks whether that relevant evidence has sufficient rigor, quality, and completeness to meet its credibility goal. Evidence may be technically strong but irrelevant, or highly relevant but inadequate, and both questions must be answered. Each goal is then recorded as met, partly met, does not meet, or not applicable with justification, rather than averaged into one score.

Some credibility factors may be designated non-compensatory for a context of use. When an unmet goal concerns such a factor, stronger evidence elsewhere does not by itself close the gap. Not every weak factor is non-compensatory; the designation is a judgement made for the specific decision.

10.4 The acceptance margin

A separate quantitative check compares the prediction against the limit. The normalized margin is

RM = (L − Sapp) / upred

the gap in units of the combined prediction uncertainty from Module 9. A decision-specific criterion such as RM ≥ 2, set from the consequence of error before the comparison, records how much margin the decision demands; equivalently L − Sapp ≥ Uallow with Uallow = 2upred. Report whether upred is a standard uncertainty or an expanded uncertainty U = kcovupred, because the criterion must be read on the same basis; 2upred is not universally a 95% coverage interval. A margin that ignores the uncertainty is meaningless, and meeting the criterion satisfies one check, not the whole credibility case.

10.5 The credibility record and the decision

The final product is a credibility record: the context of use, the question and quantity of interest, model influence, decision consequence, and the model-risk classification; the factor-specific goals with each judged for relevance and adequacy; the unmet goals and any non-compensatory gaps; the prediction uncertainty and its basis; the decision-specific margin criterion and result; the residual limitations and restrictions on use; the remediation options for unmet goals; the configuration and software version; and the conditions that would require reassessment.

When goals are not met, the response is not automatically "collect more validation data". Depending on the gap, the options include additional code or solution verification, better uncertainty characterization, more relevant validation evidence, narrowing the context of use, reducing the model's influence by adding independent evidence, introducing conservative restrictions, revising the decision criterion, or deferring the decision.

The credibility assessment is configuration- and context-specific. Changes to geometry, material, solver, mesh, calibration, boundary conditions, operating regime, software version, surrogate, output quantity, decision threshold, or context of use require impact assessment and may require renewed V&V. And the record informs the decision owner: it documents what the evidence supports, what remains uncertain, and what restrictions apply; it does not make or own the engineering, organizational, or regulatory decision.

Engineering connection: a digital twin or simulation replacing a physical test can only do so when its credibility record shows the evidence matches the decision's risk, the bridge to Digital Engineering Foundations.

04

Worked example 1: judging model risk

A simulation predicts the peak stress in a specified load-bearing implant under a specified loading case, to decide whether the part is fit to field. The question of interest is whether the peak stress stays within the allowable limit; the quantity of interest is peak stress, and the model output is the predicted peak stress. No bench test covers this loading case, so the model is the primary evidence for the decision. If the prediction is wrong, the part can fail in service. Judge the model risk and the VVUQ rigor it demands.

Figure 1. A qualitative 3×3 model-risk matrix: model influence across, decision consequence up, each judged Low, Moderate, or High. The categories are read, not multiplied or numerically spaced. This use sits in the High influence, High consequence cell, the region demanding the fullest VVUQ evidence.
  1. ProblemJudge the model risk for the context of use in Figure 1, and the VVUQ rigor it demands.
  2. Given / findOne stated context of use: peak stress in a load-bearing implant, no bench test for this loading case. Find the level of model influence, the level of decision consequence, and the rigor implied.
  3. AssumptionsRisk is judged against one stated context of use. Change the context and the judgement changes. Influence and consequence are assessed as levels, not measured.
  4. ModelModel risk rises as model influence rises and as decision consequence rises. There is no formula: the two are judged on their own axes and read together.
  5. Model influenceNo bench test covers this loading case, so the simulation carries most of the evidence for the decision rather than supporting a test that already answers it. Influence is high.
  6. Decision consequenceA wrong prediction allows a load-bearing part to fail in service. Consequence is high.
  7. Read togetherHigh on both axes is the region demanding the fullest evidence: code and solution verification, validation against relevant data, and quantified uncertainty reported with the prediction.
  8. CheckMove one axis and test the answer. Add a bench test that carries most of the decision and influence falls, so lighter model evidence can be defensible. Notice the consequence has not changed: the part can still fail. That is why both axes are needed, and why a single number would hide which one moved.
  9. ConclusionHigh influence with high consequence means the full VVUQ evidence chain. State the context of use alongside the judgement, because the same model in a different decision can carry very different risk.
Result. High influence and high consequence: full VVUQ evidence required.
05

Worked example 2: the margin criterion

A design limit is L = 100 (normalised). The simulation predicts Sapp = 85 with a combined prediction standard uncertainty upred = 6, on the same basis as the criterion. A decision-specific criterion of RM ≥ 2, set before the comparison, applies. Find the normalized margin and say whether the criterion is met.

Figure 2. Prediction, limit, and uncertainty on one scale (18 px per unit). The gap L − Sapp = 15 is 2.5 standard-uncertainty units; the dashed band is one standard prediction uncertainty (±6), and the required allowance is 2upred = 12, which the gap exceeds.
  1. ProblemFind the normalized margin RM in Figure 2 and say whether the criterion RM ≥ 2 is met.
  2. Given / findL = 100, Sapp = 85, upred = 6 as a combined standard uncertainty, criterion RM ≥ 2. Find RM.
  3. Assumptionsupred = 6 is the combined prediction standard uncertainty from Module 9, on the same basis as the criterion; the limit is treated as firm; the required margin of 2 is a decision-specific criterion set from the consequence of error, not a universal value.
  4. ModelRM = (L − Sapp)/upred, compared against the decision-specific criterion. Equivalently, L − Sapp ≥ Uallow with Uallow = 2upred.
  5. EquationsRM = (L − Sapp)/upred
  6. SolveRM = (100 − 85)/6 = 15/6 = 2.5, so the prediction sits 2.5 standard-uncertainty units below the limit. Since 2.5 ≥ 2, the margin criterion is met on the declared standard-uncertainty basis.
  7. CheckHad the uncertainty been 8 instead of 6, RM = 15/8 = 1.9, below the criterion, showing how the uncertainty, not just the gap, decides the margin. Had upred instead been reported as an expanded uncertainty, the criterion would have to be restated on that basis.
  8. ConclusionRM = 2.5 exceeds the required 2, so the stated margin criterion is met on the declared basis. This is one item in the credibility record; design acceptance still depends on whether the full set of credibility goals is met, whether the residual limitations are acceptable to the decision owner, and whether the other required evidence is satisfied.
Result. RM = 2.5, so the stated margin criterion RM ≥ 2 is met on the declared standard-uncertainty basis. It is one input to the decision, not the decision.
06

Misconceptions and diagnostics

MistakeSymptomDiagnostic questionCorrection
Uniform rigorSame evidence for every model"What is the model risk?"Scale VVUQ evidence to influence and consequence.
Credibility without a context of useA model called valid in general"Credible for what decision?"Judge credibility for a stated context of use.
Relevant confused with adequateStrong evidence for the wrong phenomenon"Does this challenge what matters here?"Judge relevance and adequacy separately; a critical gap is not offset elsewhere.
Goals set after the resultsThe bar lowered to fit the evidence"Was this goal set before the evidence?"Set factor-specific goals from the risk, before assessing evidence.
Margin criterion read as acceptanceA design accepted on RM alone"Are the other goals met, on what basis?"The margin is one item; acceptance needs the full record and the decision owner.
Margin without uncertainty basisGap reported as if exact, or a standard-deviation label with no stated distribution"RM on standard or expanded uncertainty?"Divide the gap by upred and state its basis; 2upred is not automatically 95%.
Analyst owns the decisionCredibility record used as the approval"Who is accountable, and what must be reassessed?"The record informs the decision owner; changes to model or context require reassessment.
07

Practice ladder

Level 1 · Direct skill

A model informs a design choice that a physical test will also check, and a wrong choice would be caught long before anything is built. Judge the model risk and the evidence it calls for.

Show answer

Model influence is low: the test carries most of the decision. Decision consequence is low: an error is caught early. Low on both axes is the region calling for the least evidence, so basic verification and a sanity comparison can be enough. Say which axis you judged low and why, because that reasoning is the answer, not a number.

Level 2 · Mixed concept

A limit is 50, the prediction is 40, and the combined prediction standard uncertainty is 4. Find the normalized margin RM.

Show answer

RM = (50 − 40)/4 = 2.5, so the prediction is 2.5 standard-uncertainty units below the limit, on the stated basis.

Level 3 · Independent problem

For a criterion of RM ≥ 3, a prediction of 70 against a limit of 100 needs what maximum prediction uncertainty?

Show answer

RM = (100 − 70)/u ≥ 3 ⇒ 30/u ≥ 3 ⇒ u ≤ 10. The prediction uncertainty must be 10 or less to meet the criterion, read on the stated basis.

Transfer task | Real engineering

You want a simulation to replace a physical qualification test. Describe the credibility case you would build.

What good work looks like

State the context of use and the question and quantity of interest; judge the model risk from its influence and the consequence of failure; set factor-specific credibility goals from that risk before assessing evidence; assemble verification, validation, and quantified uncertainty, judging each for relevance and adequacy; show the normalized margin against the limit on a stated uncertainty basis; and document the unmet goals, residual limitations, restrictions on use, what would require reassessment, and who owns the decision. If a goal is unmet, name the remediation rather than defaulting to more validation data.

08

Working with AI, and proving it yourself

Use AI as an examiner, not a solver

"Check that my required rigor matches the model risk, and flag any goal I set after seeing results."
"Give me three cases; I will compute RM and state its uncertainty basis for each."
"Check whether each item of my evidence is relevant to the context of use, not just strong."
"Decide if my model is credible" or "assign the influence, consequence, or goals." These are stakeholder judgements.
"Decide the acceptable residual risk" or "confirm regulatory compliance." The decision owner and the regulator hold these, not the model or the AI.

Portfolio task

For a real simulation-based decision, state the context of use, judge the model risk, set factor-specific goals, assemble the VVUQ evidence and judge each item for relevance and adequacy, compute the normalized margin on a stated uncertainty basis, and write a short credibility record.

Must include: a stated context of use, a reasoned risk judgement, factor-specific goals with relevance and adequacy noted, a normalized margin with its uncertainty basis, residual limitations, and the decision owner.
09

Retrieval and spaced review

Closed notes. Answer out loud, then reveal.

1. What is credibility?

VVUQ evidence that is both relevant and adequate for a specific context of use, judged against goals set from the model risk.

2. What determines model risk?

The model's influence on the decision and the consequence of being wrong, judged as levels and read together, not multiplied and not a probability.

3. What is the difference between relevance and adequacy?

Relevance asks whether the evidence challenges what matters to the context of use; adequacy asks whether it is rigorous and complete enough to meet its goal. Both are required.

4. Write the normalized margin.

RM = (L − Sapp)/upred, in standard-uncertainty units, compared to a decision-specific criterion on a stated basis.

5. What does the credibility record not do?

It informs the decision owner and states residual limitations and what would require reassessment; it does not make or own the engineering, organizational, or regulatory decision.

TodayFinish this quiz and Levels 1 and 2 of the ladder.
+1 dayRe-derive the risk judgement and the margin from a blank page.
+3 daysAssess the credibility of three simulation uses.
+7 daysCombine verification, validation, and UQ into one credibility case.
+30 daysCarry VVUQ into Digital Engineering Foundations, the recommended next course, and revisit the course through the VVUQ hub.