Confidence intervals for simulation results

Show the statistical uncertainty around a result from repeated simulation runs.

On this page
  1. Report uncertainty around the estimate
  2. Do not read a confidence interval as a range of future runs
  3. Narrower intervals need more information
  4. For A/B decisions, estimate the difference directly

Report uncertainty around the estimate

An average from repeated runs is still an estimate. A confidence interval shows the statistical uncertainty in that estimate under the assumptions of the analysis.

For independent replications, a common interval for a mean uses the sample mean, sample standard deviation, number of replications, and a Student-t critical value.

mean ± t × sample standard deviation / √replications

Advanced details: when this interval is valid

The usual Student-t interval treats the replication-level estimates as independent observations. Raw observations collected close together inside one steady-state run are usually serially correlated and should not be substituted for independent replications.

For paired comparisons, form one difference per matched replication and build the interval from those differences rather than treating the two alternatives as unrelated samples.

Do not read a confidence interval as a range of future runs

A 95% confidence interval for mean waiting time is about the unknown mean, not about where 95% of individual patients or future replications will fall.

If you need a tail risk such as the 95th percentile of patient waiting time, estimate that quantity directly and report its uncertainty separately.

Narrower intervals need more information

More independent replications usually narrow the interval. High run-to-run variation widens it. The useful question is whether the interval is narrow enough to distinguish decisions that matter.

For A/B decisions, estimate the difference directly

If the decision is whether option B improves on option A, calculate B minus A for matched replications when the experiment design supports pairing. An interval on the difference answers the decision more directly than two separate intervals.

References