Queues, capacity, and bottlenecks

Learn why waiting forms, how limited resources shape flow, and how to identify the constraint that actually limits performance.

On this page
  1. Why queues form
  2. Resources set the available capacity
  3. Read utilization in context
  4. A bottleneck limits the whole system

Why queues form

A queue forms when work arrives faster than the next step can accept it. This can happen briefly even when average capacity exceeds average demand, because arrivals and service times vary.

Queue length tells you how much work is waiting. Waiting time tells you how long it waits. Both depend on the surrounding process, not on the queue alone.

Resources set the available capacity

Resources are limited things needed to perform work: people, machines, beds, rooms, tools, or vehicles. Capacity is how many units of work can be handled at once.

Calendars, breaks, failures, and competing tasks reduce effective capacity. A model should include them when they can change where or when waiting occurs.

Read utilization in context

Utilization is the share of available capacity that is busy. High utilization may look efficient, but a resource operating near 100% has little room to absorb a burst of arrivals or a long job.

Never read utilization by itself. Compare it with waiting time, queue length, throughput, and service level. Use the example below to see how these measures move together.

A bottleneck limits the whole system

The busiest resource is not necessarily the bottleneck. A bottleneck is the constraint whose capacity changes a system-wide result such as throughput or total delay.

Test a suspected bottleneck by changing its capacity while holding the rest of the model constant. If performance barely changes, the real constraint is elsewhere. After one constraint is relieved, another may take its place.