How to run a warehouse automation pilot
A practical guide to structuring, running, and evaluating an automation pilot that generates the evidence needed to make a full deployment decision with confidence.
8 min read · Vendor-agnostic
Why pilots are often poorly structured
Many automation pilots are treated as demos rather than as structured evaluations. The technology is deployed, staff observe it working, and the decision to proceed is made on impressions rather than evidence. This approach produces commitment without validation.
A well-structured pilot answers specific questions about your operation, generates documented KPI evidence, and produces a formal review that justifies the full deployment decision. The goal is not to see the technology working in general, but to prove it works in your specific environment at the performance levels your business case assumes.
What a pilot must answer
Before the pilot begins, define the specific questions it must answer. These become the evaluation criteria against which the pilot will be formally assessed.
Does the technology achieve the required throughput?
Not under ideal conditions, but under your actual operational conditions: shift patterns, peak loads, WMS integration latency, and real exception rates. Measure over at least four weeks, not just during the commissioning period.
Is WMS integration stable?
How many integration errors occur per day? What is the average resolution time? Does the integration perform consistently across peak and off-peak periods? Is the exception handling process manageable for operations staff?
What is the actual exception rate?
How often does the system require manual intervention? What causes exceptions? Is the exception handling process adding labour rather than reducing it? Are exceptions declining as the system matures?
What is the vendor support quality?
How responsive is vendor support during incidents? What is the actual resolution time for the issues that arose during the pilot? Does the support model function as described in the contract?
Are the business case assumptions holding?
Compare actual pilot performance against the assumptions in your business case. Labour saving rate, throughput, system uptime, and exception rate should all be tracked against the projected values.
Pilot KPIs
Record all KPIs before the pilot begins (baseline) and track them weekly during the pilot. KPIs without baselines cannot demonstrate improvement.
| KPI | What to measure |
|---|---|
| System uptime | Percentage of scheduled operating time the system was available |
| Throughput rate | Tasks or movements completed per hour under operational conditions |
| Exception rate | Percentage of tasks requiring manual intervention |
| WMS integration errors | Number of integration errors per day and average resolution time |
| Labour utilisation | FTE required in the piloted area versus pre-pilot baseline |
| Order accuracy | Error rate per 1,000 orders in the piloted area |
| Vendor support response time | Time from incident report to first vendor response |
| Vendor resolution time | Time from incident report to system restoration |
Pilot scope and duration
A pilot should be large enough to generate statistically meaningful data but small enough to contain risk. Define scope and duration before the pilot begins.
Recommended pilot duration
- Minimum 8 weeks of operational data after full commissioning
- Include at least one peak period if your operation has seasonal volume variation
- Do not use the first 2 weeks of commissioning data as pilot evidence (stabilisation period)
- Plan for a formal review meeting at week 4 and a go/no-go decision at week 8
For AMR pilots, a minimum of 5-10 robots in a defined zone provides enough operational data. For GTP picking systems, pilot a minimum of two pick stations and one complete product category before committing to full deployment.
Common pilot traps
Using commissioning performance as pilot evidence
The first weeks after deployment are the best-case performance period: the environment is clean, staff are engaged, and the vendor is on-site. Commissioning performance is not a reliable predictor of steady-state performance.
No baseline recorded before go-live
Without a pre-pilot baseline, you cannot measure improvement. Record manual throughput, labour utilisation, and accuracy in the target area before the technology is deployed.
Vendor-defined success criteria
If the vendor defines the KPIs and the success thresholds, the pilot is structured to pass. Define your own success criteria before vendor engagement.
Pilot designed to succeed, not to test
Pilots that run only under ideal conditions, avoid peak periods, or use a more favourable product mix than normal operations will not generate reliable evidence for the full deployment decision.
No formal review with documented output
A pilot without a formal review report and a documented go/no-go decision leaves the deployment authorisation process without a defensible evidence base.
Scaling pressure before review is complete
Vendors often apply commercial pressure to expand scope before the formal pilot review is complete. Resist this pressure until the review has been formally closed with documented KPI evidence.
From pilot to full deployment
The formal pilot review should produce a documented decision: proceed to full deployment, extend the pilot to address specific gaps, or exit the programme with the vendor.
If the pilot showed performance below business case assumptions, the review should identify whether the gap is recoverable (operational improvement, configuration change, additional training) or structural (the technology does not fit the operation as designed).
A successful pilot review should update the business case with actual pilot performance data, replacing the original planning assumptions with evidence. This revised business case is the basis for the full deployment investment approval.
Plan your automation programme
Use the workspace to track your pilot KPIs, document findings, and build the deployment case from real data.
Start Free Assessment