Simulation input modelling from data
Choose how service times, arrivals, failures, and other random inputs should be represented from observed data.
On this page
- Start with the process and the data you actually have
- An empirical distribution can be the simplest honest choice
- Fit a named distribution when it helps
- Preserve dependence that changes system behavior
- Understand how the dataset was filtered and measured
- Do not force one distribution onto a changing process
- With little data, expose the assumption instead of inventing precision
- A fitted curve is still an estimate
Input modelling decides how random inputs such as service times, repair times, demand, or interarrival times are represented in the simulation.
Do not choose a distribution because its shape looks familiar. Check how the data were collected, whether the process changed during the sample, and whether observations are independent enough for the planned model.
If you have enough representative observations, sampling from the observed distribution can preserve features that a fitted textbook distribution would smooth away.
The tradeoff is that an empirical distribution contains only what was observed. It may be poor at representing rare tails or future conditions outside the sample.
A named distribution such as gamma, log-normal, or Weibull can make scenario changes and tail assumptions easier to state. Compare candidate distributions with the data and with what is physically possible, including minimums, maximums, and rare long values.
Input observations are not always independent. A difficult job may require both longer service and more rework; demand may be correlated across products; successive interarrival times may cluster during busy periods.
Sampling each field independently can destroy those relationships and understate congestion or risk. When observations belong together, preserve the pair, vector, time series, or cycle rather than fitting each column in isolation.
Ask what is missing from the data. Completed-service records may exclude customers who abandoned the queue; repair logs may omit minor failures; rounded timestamps can create artificial spikes at whole minutes.
Censoring, truncation, aggregation, and selection rules can matter more than the choice between two similar fitted distributions. Document those rules before using the sample as a model of the real process.
If the process changes by hour, weekday, season, product type, or operating regime, split or condition the input model where that variation is real. Mixing several regimes into one distribution can produce a shape that matches none of them.
For repeated time patterns, preserve complete cycles when possible rather than independently resampling individual observations and losing the time structure.
When measurements are sparse, use engineering limits, process constraints, expert estimates, or comparable systems to define a plausible range. A triangular or bounded distribution can be more honest than fitting several decimal places from five observations.
Then vary the uncertain assumption in sensitivity analysis. If the decision changes across plausible inputs, collect better data before treating the result as settled.
A finite dataset does not reveal the true input distribution exactly. If the decision changes when you use another plausible distribution, show that sensitivity instead of presenting one fitted curve as certain.