Utilization, capacity, and waiting time

Relate demand to service capacity and understand why waiting rises before a system reaches 100% average load.

On this page
  1. Start with demand and service capacity
  2. A system can queue before average demand reaches average capacity
  3. Separate current occupancy from utilization over time
  4. Increasing the wrong capacity will not fix the queue

Start with demand and service capacity

For identical parallel servers, a rough offered-load calculation compares the arrival rate with total mean service capacity. If one server completes μ jobs per hour on average and there are c servers, total mean capacity is cμ.

offered load ρ = λ / (cμ)

A system can queue before average demand reaches average capacity

Arrivals do not come at perfectly even intervals and jobs do not all take the same time. Bursts and long jobs create temporary backlogs, so waiting can become large as spare capacity shrinks.

That is why an average load below 100% does not mean nobody waits.

Separate current occupancy from utilization over time

Two of three servers busy right now means current occupancy is 66.7%. Utilization is normally a time-based measure over an interval. Do not label a single instant as if it were a long-run utilization estimate.

When diagnosing a queue, look at waiting time, queue length, throughput, and time-based utilization together.

Increasing the wrong capacity will not fix the queue

Check every stage in the path. If an upstream process can only supply 9 jobs per hour, increasing a downstream station from 11 to 15 jobs per hour will not increase total throughput.

References