Simulation input modelling from data

Choose how service times, arrivals, failures, and other random inputs should be represented from observed data.

On this page
  1. Start with the process and the data you actually have
  2. An empirical distribution can be the simplest honest choice
  3. Fit a named distribution when it helps
  4. Preserve dependence that changes system behavior
  5. Understand how the dataset was filtered and measured
  6. Do not force one distribution onto a changing process
  7. With little data, expose the assumption instead of inventing precision
  8. A fitted curve is still an estimate

Start with the process and the data you actually have

Input modelling decides how random inputs such as service times, repair times, demand, or interarrival times are represented in the simulation.

Do not choose a distribution because its shape looks familiar. Check how the data were collected, whether the process changed during the sample, and whether observations are independent enough for the planned model.

An empirical distribution can be the simplest honest choice

If you have enough representative observations, sampling from the observed distribution can preserve features that a fitted textbook distribution would smooth away.

The tradeoff is that an empirical distribution contains only what was observed. It may be poor at representing rare tails or future conditions outside the sample.

Fit a named distribution when it helps

A named distribution such as gamma, log-normal, or Weibull can make scenario changes and tail assumptions easier to state. Compare candidate distributions with the data and with what is physically possible, including minimums, maximums, and rare long values.

Preserve dependence that changes system behavior

Input observations are not always independent. A difficult job may require both longer service and more rework; demand may be correlated across products; successive interarrival times may cluster during busy periods.

Sampling each field independently can destroy those relationships and understate congestion or risk. When observations belong together, preserve the pair, vector, time series, or cycle rather than fitting each column in isolation.

Understand how the dataset was filtered and measured

Ask what is missing from the data. Completed-service records may exclude customers who abandoned the queue; repair logs may omit minor failures; rounded timestamps can create artificial spikes at whole minutes.

Censoring, truncation, aggregation, and selection rules can matter more than the choice between two similar fitted distributions. Document those rules before using the sample as a model of the real process.

Do not force one distribution onto a changing process

If the process changes by hour, weekday, season, product type, or operating regime, split or condition the input model where that variation is real. Mixing several regimes into one distribution can produce a shape that matches none of them.

For repeated time patterns, preserve complete cycles when possible rather than independently resampling individual observations and losing the time structure.

With little data, expose the assumption instead of inventing precision

When measurements are sparse, use engineering limits, process constraints, expert estimates, or comparable systems to define a plausible range. A triangular or bounded distribution can be more honest than fitting several decimal places from five observations.

Then vary the uncertain assumption in sensitivity analysis. If the decision changes across plausible inputs, collect better data before treating the result as settled.

A fitted curve is still an estimate

A finite dataset does not reveal the true input distribution exactly. If the decision changes when you use another plausible distribution, show that sensitivity instead of presenting one fitted curve as certain.

References