At a glance

What this page covers

For
Power-electronics engineers deciding whether a fast surrogate may influence design exploration.
You will leave with
A six-check acceptance packet and a four-action refusal rule for evaluating a surrogate.
Evidence status
Working acceptance framework; it reports no project-specific qualification result.
Boundary
Thresholds depend on the decision, evaluated domain, and consequence of error; this note supplies no universal acceptance value.
Starting point, decision and reusable output
Decision
Choose whether to use an estimate, show a warning, request higher-fidelity evidence, or abstain.
Starting point
No model-building experience is required; begin with the engineering decision and the reference method you already trust.
Reusable output
A reusable minimum-evidence table and a use, warn, request-higher-fidelity, abstain decision rule.
Next useful action

Apply grouped validation and a trivial reference model in the runnable beginner lab before evaluating a more complex surrogate.

Run the beginner lab

A surrogate model is a fast approximation of a slower calculation, simulation, or measurement. It can reduce waiting during design exploration, but a low average error does not tell an engineer where the approximation is safe to use. A model may perform well on familiar examples and still fail near a physical limit, on a new component family, or at the point that determines the design decision. The practical question is therefore not, “Is the model accurate?” It is, “For this design and this decision, should the workflow use the estimate, show a warning, request stronger evidence, or refuse to predict?”

Why one error number is not enough

An average combines easy and difficult cases. A model may look good overall while making its largest errors in one component family, near a material limit, or in the operating region that matters most.

This is familiar engineering behaviour. A converter is not accepted from one nominal operating point, and a surrogate should not be accepted from one summary metric.

The hidden risk in a fast, precise-looking answer

If the model is used outside the evidence that was tested, it can produce a precise-looking value with no trustworthy basis. The workflow then saves calculation time but increases decision risk.

The most dangerous case is not an obviously broken prediction. It is a plausible number that passes through the design process without anyone noticing that the input was unfamiliar.

Treat trust as permission with conditions

Trust should be written as an operating rule:

  • use the estimate inside a tested region when all checks pass;
  • warn when evidence is weaker but the result may still support exploration;
  • request a higher-fidelity method when the decision is consequential;
  • abstain when the input or output is outside the accepted contract.
Decision flow for accepting or rejecting a surrogate-model estimate
A surrogate earns limited permission through baselines, unseen-design tests, failure analysis, and a safe action. It never approves the component by itself.

Six checks in plain engineering language

1. Compare with the method you would otherwise use

Start with a trivial baseline and the best practical engineering approximation already available. If the surrogate adds little, the simpler method may remain the better system because it is easier to inspect and maintain.

2. Hold back complete designs, not random rows

Measurements from the same physical design are related. If rows from one design appear in both training and test data, the test can look easier than the real task. Group the split by component or design family so the evaluation represents genuinely unseen hardware.

3. Plot the errors against physical inputs

Do not stop at mean error. Plot the residual, which is prediction minus reference, against each meaningful input, the target value, the operating regime, and the component family.

4. Detect unfamiliar inputs

An unfamiliar-input score is useful only when it changes behaviour. Define what the workflow does when a new point is far from the evaluated data: warn, request a reference calculation, or abstain.

5. Check whether uncertainty statements are honest

If the model returns an interval, test how often the reference value falls inside that interval on grouped, unseen designs.

coverage = P(reference\ value\ is\ inside\ the\ predicted\ interval)

Coverage is not a guarantee for one component. It is a check that the stated uncertainty behaves as claimed across the evaluation set.

6. Apply physical checks and test the refusal path

Use known bounds, impossible states, conservation relationships, and any justified monotonic behaviour as independent checks. Then test the complete refusal path. A safe abstention must lead somewhere useful, such as a higher-fidelity calculation, a measurement request, or engineering review.

The minimum acceptance packet

Evidence itemQuestion it must answerDecision it supports
Versioned data and provenanceWhat designs, conditions, units, and sources were evaluated?Is this input inside the known scope?
Practical baselineDoes the surrogate improve on the method already available?Is the added complexity justified?
Grouped unseen-design resultsDoes performance survive on hardware families withheld from training?Can the estimate generalise within the stated domain?
Residual and worst-case plotsWhere are the large or structured errors?Where should the workflow warn or abstain?
Unfamiliar-input and uncertainty checksDoes confidence change when evidence becomes weak?Can the system choose a safer action?
Physical validationAre outputs plausible and independently checked?Is the result suitable for the stated engineering use?

Why no universal confidence threshold exists

The acceptable error and confidence policy depend on the decision. Early design-space exploration can tolerate a different error from a thermal limit, a saturation check, or a component release decision.

Write the refusal rule before training

Before choosing a model, write four outputs on one page: predict, warn, request stronger evidence, and abstain. For each output, state the evidence required and who owns the next decision.

That small contract prevents a good-looking metric from becoming accidental engineering authority. Subscribe to Field Notes for the next worked case, where the same evidence-first approach is applied to power-electronics data and automation.