How long should a simulation run?

Choose a simulation horizon from the study type, system time scale, rare events, and the precision needed for the decision.

On this page
  1. Run long enough for the output you need to estimate
  2. For terminating studies, use the natural end condition
  3. For steady-state studies, separate warm-up from collection
  4. Rare events can dominate the required horizon
  5. Stop when the estimate is precise enough for the decision

Run long enough for the output you need to estimate

Run length is not a software default. It follows from the study type, the time scale of the system, and the measure you need to estimate.

A horizon that is adequate for hourly throughput may be useless for annual failures or a 99th-percentile delay. Choose the horizon against the slowest important behavior in the decision.

For terminating studies, use the natural end condition

If the real episode ends after one operating day, after all scheduled jobs finish, or when an evacuation completes, let that event define the replication. Extending the simulation simply to collect more observations changes the question.

Need more precision? Repeat the full episode with independent random streams rather than quietly turning one day into an artificial multi-day study.

For steady-state studies, separate warm-up from collection

After the warm-up period, the remaining run must still be long enough to represent normal variation. Nearby observations inside one long run are often correlated, so a million recorded time points are not a million independent observations.

Independent replications, replication/deletion, or batch-means methods are common ways to obtain usable uncertainty estimates from steady-state output.

Advanced details: batch means

Batch means divides one long post-warm-up run into consecutive blocks and analyzes the block averages. The batches must be long enough that dependence between neighboring batch means is small enough for the intended interval calculation.

Do not manufacture a large sample size by treating every timestamp or entity observation from one correlated run as independent.

Rare events can dominate the required horizon

A reliability model that fails once every few months cannot support a failure-rate conclusion from a two-day run. The same problem appears with severe congestion, stockouts, tail waiting times, and other rare outcomes.

If most replications contain zero occurrences of the event you care about, increase the exposure or use a method designed for rare-event analysis instead of reporting zero as evidence of safety.

Stop when the estimate is precise enough for the decision

Use a pilot run to learn the time scale and variability of the output. Then choose run length and replication count so the uncertainty is small relative to the difference that would change the decision.

Longer replications and more replications solve different problems. Longer runs reveal slow behavior and reduce within-run noise; more independent replications give more independent estimates for uncertainty analysis.

References