A Pilot That Cannot Fail Is a Subscription
Ninety days later nobody can say whether it worked, because nobody agreed a number in advance. The four things to fix before a single system is connected, and the sentence a vendor should be willing to put in writing.
See how Ward detects pilots with pre-agreed close criteria
Get a demo → Take the 3-minute assessmentContents
- Ninety days later, nobody can say whether it worked
- Everything is decided before anything connects
- One KPI, not a scorecard
- A baseline you both signed
- Agreement on what counts as attribution
- An explicit stop condition
- What the ninety days should look like
- Publishing the result, including a bad one
- Why this is in your interest even when the tool is good
Ninety days later, nobody can say whether it worked
The pattern is familiar enough to be boring. A retailer runs a three-month pilot of an analytics or AI tool. At the end there is a readout. The readout contains engagement metrics, a few anecdotes from store managers, a slide of screenshots, and a recommendation to expand.
What it does not contain is a number that anyone agreed to in advance.
So the decision gets made on vibes and internal politics: whoever sponsored it wants to continue, whoever did not wants to stop, and the tool's actual effect on the business is unmeasured in either direction. The vendor is happy, because ambiguity favours the incumbent. The retailer has spent a quarter and learned nothing transferable.
This is avoidable, and it is avoidable entirely in the first week.
Everything is decided before anything connects
Close criteria are worth nothing if they are written after you have seen the data. Four things have to be fixed before a single system is connected.
One KPI, not a scorecard
Pick the metric and the threshold it has to clear. One. A scorecard of six measures is a way of guaranteeing that something moved, which is the same as guaranteeing nothing is falsifiable.
The metric should be one your finance team already reports, not one the tool introduces. If the only way to see the improvement is through the vendor's own instrumentation, you have not measured the business, you have measured the tool.
A baseline you both signed
Fix the pre-pilot number from your own history and freeze it. State the comparison window in advance: which weeks, which stores, adjusted for what.
This is the step that gets skipped and the one that matters most. If the baseline is chosen after the result is known, every pilot succeeds, because there is always a window in which the number looks better. Neither side should be able to pick the comparison retrospectively, including you.
Agreement on what counts as attribution
Retail KPIs move for reasons that have nothing to do with software. Weather, a competitor closing, a supply disruption, a category reset that was already planned. Decide up front how you will separate those from the intervention, and write down what would make you say "this moved for another reason."
Control stores are the cheapest version of this. If the pilot runs in 40 stores and the other 300 provide the counterfactual, most confounders take care of themselves.
An explicit stop condition
The sentence that has to appear: if the KPI has not moved by day 90, the pilot ends. Not "we will review", not "we will consider a second phase." Ends.
A vendor who will not put that in writing is telling you they expect the ambiguity to work in their favour, which it will.
See how Ward detects pilots with pre-agreed close criteria
Get a demo →What the ninety days should look like
The sequence matters less than the fact that it is agreed, but a shape that works in retail looks roughly like this.
First 48 hours. Read-only connection and first findings on your own data. Not a sandbox, not a demo environment. If a tool cannot say anything useful about your actual data in two days, the problem is unlikely to be time.
Week 2. Findings ranked by recoverable dollars, with a cause attributed to each. The useful test here is not whether the findings are impressive; it is whether your operators already knew them. A tool that surfaces only things your best district manager could have told you is working, but it is not worth a platform decision.
Week 6. The first action taken under approval. This is the real gate: everything before it is analysis, and analysis has never been the bottleneck in retail. What matters is whether the loop closes into the system of record with a named human approving each write.
Day 90. The measured delta against the baseline you froze in week zero, published whichever way it went.
Publishing the result, including a bad one
Ask the vendor, before signature, what they will publish if the pilot fails. Most have not been asked and it shows.
The answer worth having is specific: the KPI delta in both directions, win rate per playbook including the ones that did not close, an explicit statement where a metric moved for reasons other than the tool, and whether the customer renewed. Committing to that shape in advance is what stops the report being written backwards from the numbers.
We hold ourselves to it publicly, and our own result is still unpublished: a production pilot inside an enterprise grocery estate, running against their live POS, ERP and inventory, with no measured result yet. The pilot status page says exactly that rather than filling the gap with a projection, because the projection lives elsewhere labelled as modelled.
Why this is in your interest even when the tool is good
It is tempting to read all of this as protection against a bad vendor. It is more useful as protection against a good one.
A pilot with real close criteria tells you the size of the effect, not just its direction. That number is what you need to decide whether to roll out across the estate or stop at the pilot stores, what to budget next year, and what to tell a board that will ask why the line did not move as much as the business case said.
Without it you get a renewal decision made on enthusiasm, which is the same decision you would have made without running the pilot at all.
See how Ward detects pilots with pre-agreed close criteria
Ward monitors your stores 24/7 and delivers insight cards, not dashboards. First cards in 48 hours.