At a glance
What this page covers
- For
- Power-electronics engineers deciding whether a fast surrogate may influence design exploration.
- You will leave with
- A six-check acceptance packet and a four-action refusal rule for evaluating a surrogate.
- Evidence status
- Working acceptance framework; it reports no project-specific qualification result.
- Boundary
- Thresholds depend on the decision, evaluated domain, and consequence of error; this note supplies no universal acceptance value.
Starting point, decision and reusable output
- Decision
- Choose whether to use an estimate, show a warning, request higher-fidelity evidence, or abstain.
- Starting point
- No model-building experience is required; begin with the engineering decision and the reference method you already trust.
- Reusable output
- A reusable minimum-evidence table and a use, warn, request-higher-fidelity, abstain decision rule.
Apply grouped validation and a trivial reference model in the runnable beginner lab before evaluating a more complex surrogate.
A surrogate model is a fast approximation of a slower calculation, simulation, or measurement. It can reduce waiting during design exploration, but a low average error does not tell an engineer where the approximation is safe to use. A model may perform well on familiar examples and still fail near a physical limit, on a new component family, or at the point that determines the design decision. The practical question is therefore not, “Is the model accurate?” It is, “For this design and this decision, should the workflow use the estimate, show a warning, request stronger evidence, or refuse to predict?”
Why one error number is not enough
An average combines easy and difficult cases. A model may look good overall while making its largest errors in one component family, near a material limit, or in the operating region that matters most.
This is familiar engineering behaviour. A converter is not accepted from one nominal operating point, and a surrogate should not be accepted from one summary metric.
The hidden risk in a fast, precise-looking answer
If the model is used outside the evidence that was tested, it can produce a precise-looking value with no trustworthy basis. The workflow then saves calculation time but increases decision risk.
The most dangerous case is not an obviously broken prediction. It is a plausible number that passes through the design process without anyone noticing that the input was unfamiliar.
Treat trust as permission with conditions
Trust should be written as an operating rule:
- use the estimate inside a tested region when all checks pass;
- warn when evidence is weaker but the result may still support exploration;
- request a higher-fidelity method when the decision is consequential;
- abstain when the input or output is outside the accepted contract.

Six checks in plain engineering language
1. Compare with the method you would otherwise use
Start with a trivial baseline and the best practical engineering approximation already available. If the surrogate adds little, the simpler method may remain the better system because it is easier to inspect and maintain.
2. Hold back complete designs, not random rows
Measurements from the same physical design are related. If rows from one design appear in both training and test data, the test can look easier than the real task. Group the split by component or design family so the evaluation represents genuinely unseen hardware.
3. Plot the errors against physical inputs
Do not stop at mean error. Plot the residual, which is prediction minus reference, against each meaningful input, the target value, the operating regime, and the component family.
4. Detect unfamiliar inputs
An unfamiliar-input score is useful only when it changes behaviour. Define what the workflow does when a new point is far from the evaluated data: warn, request a reference calculation, or abstain.
5. Check whether uncertainty statements are honest
If the model returns an interval, test how often the reference value falls inside that interval on grouped, unseen designs.
Coverage is not a guarantee for one component. It is a check that the stated uncertainty behaves as claimed across the evaluation set.
6. Apply physical checks and test the refusal path
Use known bounds, impossible states, conservation relationships, and any justified monotonic behaviour as independent checks. Then test the complete refusal path. A safe abstention must lead somewhere useful, such as a higher-fidelity calculation, a measurement request, or engineering review.
The minimum acceptance packet
| Evidence item | Question it must answer | Decision it supports |
|---|---|---|
| Versioned data and provenance | What designs, conditions, units, and sources were evaluated? | Is this input inside the known scope? |
| Practical baseline | Does the surrogate improve on the method already available? | Is the added complexity justified? |
| Grouped unseen-design results | Does performance survive on hardware families withheld from training? | Can the estimate generalise within the stated domain? |
| Residual and worst-case plots | Where are the large or structured errors? | Where should the workflow warn or abstain? |
| Unfamiliar-input and uncertainty checks | Does confidence change when evidence becomes weak? | Can the system choose a safer action? |
| Physical validation | Are outputs plausible and independently checked? | Is the result suitable for the stated engineering use? |
Why no universal confidence threshold exists
The acceptable error and confidence policy depend on the decision. Early design-space exploration can tolerate a different error from a thermal limit, a saturation check, or a component release decision.
Write the refusal rule before training
Before choosing a model, write four outputs on one page: predict, warn, request stronger evidence, and abstain. For each output, state the evidence required and who owns the next decision.
That small contract prevents a good-looking metric from becoming accidental engineering authority. Subscribe to Field Notes for the next worked case, where the same evidence-first approach is applied to power-electronics data and automation.