Warehouse KPIs That Actually Matter: What to Measure and What to Ignore
Most warehouse KPI lists are written by people selling software. This guide covers the eight metrics experienced operators actually run on, how to calculate each one honestly, what good looks like for different operation types, and the measurement mistakes that make a number worse than useless.
TL;DR: Eight metrics cover almost everything a warehouse operator needs: order accuracy, on-time shipment, dock-to-stock time, inventory accuracy, cost per order shipped, capacity utilization, units per labor hour, and perfect order rate. Each one is defined below with an honest formula, what good looks like across different operation types, and the specific way it gets gamed or mismeasured. If you track nothing today, start with order accuracy and inventory accuracy — they are the two that tell you whether anything else you measure can be trusted.
Why Most Warehouse KPI Lists Are Useless
Search for warehouse KPIs and you will find dozens of lists of twenty or thirty metrics. They share three problems.
They are written to sell dashboards. A metric that is hard to calculate by hand is a metric that needs software. Lists produced by software vendors skew heavily toward the metrics their software happens to produce, which is not the same as the metrics that should drive your decisions.
They give one benchmark for every operation. "Best in class order accuracy is 99.9%" is meaningless without knowing whether we're discussing single-line e-commerce picks or 40-line mixed-case pallet builds for grocery. The same number represents wildly different performance across operation types, and a benchmark that ignores that will either flatter you or demoralize you, at random.
They don't say how the number gets gamed. Every metric on this page can be made to look better without the operation getting better. That is the most important thing to know about each one, and it is almost never included.
What follows is eight metrics. Not twenty. Each with a formula, a realistic range by operation type, and the specific way it lies to you.
A note on the ranges below: they are drawn from operational experience, not from a published benchmark study. Treat them as orientation, not as targets handed down from an authority. See the note on benchmarks for why nobody should be quoting you precise industry figures with a straight face.
1. Order Accuracy
What it measures: the share of orders shipped without a picking, packing, or quantity error.
Formula: orders shipped correctly ÷ total orders shipped × 100
Why it comes first: it is the metric your customer feels. Everything else on this list is internal; this one shows up in their inbox as a complaint. It is also the metric most directly tied to whether you keep the account.
What reasonable performance looks like by operation type:
| Operation type | Typical range | Notes |
|---|---|---|
| E-commerce, single-line picks | 99.5–99.9% | Low line complexity; errors are usually wrong-item or wrong-quantity |
| Multi-line B2C fulfillment | 99.0–99.7% | Error probability compounds with lines per order |
| Case-pick B2B distribution | 98.5–99.5% | Fewer, larger picks; errors are more costly per occurrence |
| Mixed-case pallet building | 97.5–99.0% | Highest complexity; a single pallet may have 40+ SKUs |
How it lies to you: almost every operation measures this as reported errors, not actual errors. You are measuring customer complaints, which is a measure of how much your customer checks. A retail client with receiving-side scanning finds everything. A client who counts nothing finds nothing, and your accuracy against them looks superb right up until the annual reconciliation.
The fix: audit a sample of outbound orders yourself before they ship — even 20 a week — and track that number separately as your audited accuracy. When audited accuracy and reported accuracy diverge sharply, the gap is your customer's detection rate, not your performance.
2. On-Time Shipment
What it measures: the share of orders that shipped by the deadline they were supposed to ship by.
Formula: orders shipped on or before cutoff ÷ total orders due × 100
The definitional trap that makes this metric worthless: on time against what? Three different definitions are in common use, and they produce very different numbers:
- Against your dock cutoff — did it leave the building by your stated time
- Against the customer's requested ship date — did it leave when they asked
- Against carrier pickup — did it physically get on a truck
Pick one, write it down, and never change it silently. An operation that quietly shifts from the third definition to the first will show a dramatic improvement that never happened.
Typical range: 97–99.5% for a well-run operation on the dock-cutoff definition. Lower on the customer-requested-date definition, because that one includes orders that arrived too late to process — which is a sales and onboarding problem, not a floor problem.
How it lies to you: orders that were never accepted into the queue don't appear in the denominator. If your system rejects or holds an order and it's never counted as due, you can run 100% on-time while a stack of held orders ages in the corner. Track your held and exception queue as a separate number, always.
3. Dock-to-Stock Time
What it measures: elapsed time from a truck arriving to the inventory being putaway and available to pick.
Formula: average of (time available to pick − time of trailer arrival)
Why it matters more than operators think: dock-to-stock is the leading indicator for half the problems you'll have next week. Inventory that is on the floor but not in the system is inventory you can't sell, can't count accurately, and will trip over. Long dock-to-stock times are also the most common hidden cause of inventory inaccuracy, because product sitting in a receiving staging area gets partially putaway, partially picked from, and never reconciled.
Typical ranges by operation type:
| Operation type | Typical range |
|---|---|
| Cross-dock / flow-through | Under 4 hours |
| E-commerce fulfillment | 4–24 hours |
| Bulk / pallet-in distribution | 24–48 hours |
| Cold storage | 2–12 hours (product integrity forces the pace) |
How it lies to you: measuring from unload complete rather than trailer arrival removes detention time from the metric entirely — which is often the worst part of the process and the part you are being billed for. Measure from arrival. If you also want an unload-to-putaway number, track it separately, don't substitute it.
If your dock-to-stock is bad and you're not sure why, the receiving process is usually where it starts. We covered the failure modes in the receiving dock survival guide.
4. Inventory Accuracy
What it measures: how closely system quantities match physical quantities.
Formula (location-level, the honest one): locations counted with zero variance ÷ total locations counted × 100
Formula (piece-level, the flattering one): 1 − (absolute variance in units ÷ total units counted) × 100
Both are legitimate. They are not interchangeable, and the second will always look better — often by several percentage points — because large-quantity locations dilute small errors. An operation reporting 99.8% piece-level accuracy might be at 96% location-level. Publish which one you use.
Typical location-level ranges:
| Operation type | Typical range |
|---|---|
| Serialized / high-value | 99.5%+ |
| E-commerce, bin-level | 98–99.5% |
| Bulk pallet storage | 97–99% |
| Mixed / non-barcoded operations | 92–97% |
How it lies to you: counting the same accessible, fast-moving, well-organized locations repeatedly. Those locations are accurate because they turn — errors surface and get fixed naturally. The variance lives in slow-moving C items in the back corner that nobody has counted in eight months. A cycle count program that doesn't reach them produces a number that describes the easy half of your warehouse.
If you don't have a cycle count program at all, that is the prerequisite for this metric meaning anything. Setting one up from scratch covers the methodology, tolerances, and variance investigation.
5. Cost per Order Shipped
What it measures: fully loaded operating cost divided by orders out the door.
Formula: total operating cost for the period ÷ orders shipped in the period
What "fully loaded" has to include: direct labor, supervision, rent and occupancy, utilities, equipment lease and maintenance, packaging and consumables, WMS and technology, and an allocation of administrative overhead. If you leave out occupancy because it feels fixed, your cost per order is fiction and you will underprice.
Why there is no useful benchmark for this one: cost per order ranges from under two dollars to over twenty depending on order profile, region, wage rates, and automation. Anyone quoting you an industry average for cost per order is quoting a number with no defensible meaning. The value of this metric is entirely in its trend within your own operation and its variance between your own clients.
How it lies to you: it moves with order mix, not just with efficiency. A month where a client shifted from single-line to multi-line orders will show cost per order rising while your operation performed identically or better. Always read it alongside lines per order and units per order, or you will chase a phantom.
The most useful cut of this metric is per client, not in aggregate. That is where you find the account that is quietly unprofitable — and it pairs directly with the billing audit process, because unbilled work shows up here as cost with no matching revenue.
6. Capacity Utilization
What it measures: how much of your usable storage you are actually using.
Formula: occupied storage positions ÷ total usable storage positions × 100
Measure positions, not square feet. Square footage utilization tells you almost nothing, because it ignores cube and it counts aisle space you can't store in. Pallet positions, bin locations, or rack positions — whatever your unit is — is the number that reflects reality.
What good looks like: 85% is the number most operators should target, and higher is usually worse, not better. Above roughly 90% you lose the ability to slot efficiently, receiving backs up because there is nowhere to put things, travel distances rise, and the cost of every other metric on this list goes up. A warehouse at 97% utilization is not maximally efficient; it is one large inbound away from chaos.
Below about 70%, you are paying for space you aren't using, which is a lease conversation. Understanding what that space actually costs you is a matter of reading the lease properly — warehouse lease terms explained covers what you're really paying for.
How it lies to you: counting positions that are technically storage but practically unusable — damaged rack, positions blocked by a staging area that never moves, the aisle nobody can get a reach truck into. Audit your denominator once a year or your utilization will look artificially low and you will make a leasing decision on a bad number.
7. Labor Productivity (Units per Labor Hour)
What it measures: throughput per hour of labor paid.
Formula: units (or lines, or orders) processed ÷ total labor hours paid
Use labor hours paid, not hours worked on the task. Paid hours include breaks, training, meetings, and the twenty minutes at shift start before anyone touches product. Those hours are real cost. Measuring against task time produces a flattering number that doesn't reconcile to payroll.
Choose your unit deliberately. Units per hour rewards large-quantity picks. Lines per hour is usually the better productivity measure because it tracks the number of discrete decisions and touches. Orders per hour matters most in single-line e-commerce. Pick the one that matches your work, state it, and stick with it.
Why no benchmark table here: productivity ranges are so operation-specific that a published figure does more harm than good. A cold-storage picker in freezer PPE and a dry-goods case picker are not comparable, and neither is comparable to an e-commerce operation with pick carts and a good slotting plan. Track your own trend and your own variance between shifts and individuals — that comparison is valid and actionable, and it's free.
For wage and employment context in the sector, the Bureau of Labor Statistics publishes warehousing and storage series that are genuinely useful for understanding your labor market, if not your productivity.
How it lies to you: it goes up when quality goes down. Any picker can go faster by checking less. Never review productivity without order accuracy next to it — a productivity gain accompanied by an accuracy decline is not a gain, it is a cost transfer to your customer service team.
8. Perfect Order Rate
What it measures: the share of orders that were right in every respect — right items, right quantity, on time, undamaged, correctly documented and invoiced.
Formula: orders with zero defects of any kind ÷ total orders × 100
Why it is worth the trouble: it is the only metric on this list that reflects the customer's actual experience. An order can be 100% accurate, on time, and still fail because it arrived damaged or the paperwork was wrong. Perfect order rate catches what the component metrics miss individually.
Expect it to be lower than you'd like. Because it is multiplicative, an operation running 99% on each of five dimensions produces a perfect order rate around 95%, not 99%. That drop is not a failure — it is arithmetic, and it is why the metric is useful. It shows the compounding cost of small defects.
Typical ranges:
| Operation type | Typical range |
|---|---|
| Mature e-commerce fulfillment | 96–99% |
| B2B distribution | 93–98% |
| Complex / high-touch operations | 88–95% |
How it lies to you: by quietly dropping a dimension. If you stop counting documentation errors because they're hard to track, the rate rises and nothing improved. Define the dimensions once, in writing, and if you change them, restate the history.
A Word on Benchmarks
Every range in this guide is an operational range, not a published industry standard, and it is worth being direct about why.
There is no authoritative, independent, publicly available benchmark study for warehouse KPIs by operation type. The figures that circulate come overwhelmingly from vendor marketing, consultancy lead magnets, and survey samples that are neither disclosed nor representative. Numbers get repeated until they acquire the appearance of authority through nothing but repetition.
The Bureau of Labor Statistics publishes genuine public-domain data on warehousing employment, wages, and injury rates, which is useful for labor-market context. It does not publish order accuracy or dock-to-stock benchmarks, and nobody credible does.
What this means practically: your own trend line is more reliable than any external benchmark. A number that improved 3% quarter over quarter in your own building tells you something true. A number compared against an unsourced industry average tells you nothing, and can send you chasing a target that doesn't apply to your operation.
The Measurement Mistakes That Ruin a Good Metric
Measuring everything. Eight metrics reviewed weekly beats thirty reviewed never. If nobody can name your KPIs from memory, you don't have KPIs, you have a report.
Changing definitions silently. The single most common way a metric becomes useless. Every definition change should be documented with a date, and history should be restated or clearly broken at that point.
Reviewing metrics nobody owns. A number with no name attached is a number nobody fixes. Each metric needs one person accountable for it.
Reviewing quality and speed separately. Productivity and accuracy must be read together, always, or you optimize one at the other's expense without noticing.
Reporting averages without distribution. An average dock-to-stock of 8 hours could be a tight cluster around 8, or half at 2 and half at 14. Those are completely different operations with completely different problems. Look at the spread.
Letting the metric become the goal. Every number here can be made to look better without the operation getting better. If you tie compensation to a single metric, you should assume it will be gamed — not out of malice, but because that is what incentives do.
If You Are Starting From Nothing
Do not implement eight metrics. Do this instead:
Month one: order accuracy and inventory accuracy. These two tell you whether anything else you measure can be trusted. Inventory accuracy in particular is foundational — if system quantities are wrong, every downstream number is built on sand.
Month two: add dock-to-stock and on-time shipment. These two cover the flow through the building, inbound and outbound.
Month three: add cost per order, by client. This is where you find out which accounts are actually making money.
After that: capacity utilization, labor productivity, and perfect order rate, as the operation and your appetite allow.
You can run the first four on a spreadsheet. None of them require software you don't have. The discipline of measuring the same thing the same way every week is worth more than any tooling, and it is the part that actually fails in most operations.
What This Guide Isn't
This is not a benchmarking study — the ranges here are operational orientation, and this guide says plainly that no credible independent benchmark set exists. It is not a WMS evaluation guide; if measuring these metrics is driving you toward software, what to put in a WMS RFP covers how to run that process without letting a vendor write your requirements. And it deliberately omits financial metrics like inventory turns and carrying cost, which mostly belong to your customer rather than to you — a 3PL is measured on execution, not on the wisdom of the inventory it was handed.