Why AI Pilots Stall in Evaluation (and How to Get Yours Out)

Key takeaways

  • A pilot that cannot fail cannot succeed either. Both outcomes need to be defined before it starts.
  • Most stalled pilots have no budget owner, which means no outcome triggers a purchase.
  • Testing "AI" in general is untestable. Test one task, done by real people, on real work.
  • Two to four weeks is the useful window. Beyond that you are collecting opinions, not evidence.

The pilot went fine. Everyone agreed it was interesting. There were some promising results and some rough edges. A follow-up session was scheduled, then moved, then quietly dropped. Nine months on, the licences are still active, three people still use them, and nobody can say whether the organisation adopted AI or not.

This is not a technology failure. It is a decision architecture failure, and it has recognisable causes. Here are the six that account for nearly all of it, and how to tell which one has hold of your pilot.

Cause 1: no exit criteria

The most common by a distance. The pilot was set up to "see how it goes", which sounds sensible and is fatal, because it defines neither success nor failure. Without either, the pilot has no natural end. It just continues at declining intensity until attention moves elsewhere.

The fix is to write, before starting, exactly what result leads to what decision. Not a vague aspiration but a sentence of this shape: if the six people in the pilot report that the weekly client summary takes meaningfully less time and the quality is acceptable to their manager, we buy 40 licences in May. If they do not, we stop and write down why.

The second half matters as much as the first. Teams avoid defining failure because it feels pessimistic. Undefined failure is precisely what turns a four-week pilot into an eighteen-month one.

Diagnostic question: if you asked everyone involved in your pilot to write down, separately, what result would cause you to stop, would the answers match? If they would not, the pilot has no exit criteria regardless of what the project document says.

Cause 2: no budget owner in the room

A pilot run entirely by people who cannot approve a purchase produces, at best, a recommendation that then has to start the funding conversation from the beginning, with a different audience, several months later. By then the enthusiasm has cooled and the evidence feels stale.

The fix is structural: whoever can approve the spend must be a participant, not a recipient of a report. Not necessarily in every session, but they should see the pilot task, meet the people doing it, and agree the exit criteria personally. That way the decision at the end is a confirmation rather than a new pitch.

Cause 3: the pilot tests the wrong thing

"Let's pilot AI" is not a testable proposition. AI is a general capability, and evaluating a general capability against no particular need produces exactly the output you would expect: a set of impressions.

A testable pilot names one workflow, one team, one measure. For example: the proposal team spends roughly a day per proposal assembling a first draft from previous documents; can that become two hours without a drop in quality that the reviewing partner notices?

That question has an answer. "Is AI useful for us?" does not.

Cause 4: the wrong people are in it

Pilots tend to be staffed with volunteers, and volunteers are systematically unrepresentative. They are more interested, more tolerant of rough edges, and often less busy than the median employee. A pilot they enjoy tells you almost nothing about whether a sceptical, fully booked colleague will adopt the same tool.

Choose participants who genuinely do the task in question, including at least one person who is unenthusiastic. If the tool wins over someone with a full workload and no particular interest in technology, you have learned something real. If it only wins over the enthusiast, you have measured enthusiasm.

Cause 5: nobody configured anything

Many pilots run on consumer accounts with default settings, no shared prompts, no access to the documents that would make the tool useful, and no guidance on which model to use. The results are correspondingly mediocre, and the organisation concludes the technology is not ready.

What it has actually tested is an unconfigured tool used by untrained people on unfamiliar tasks. That is a fair test of nothing. Even a short pilot deserves the right tier, a handful of worked examples, and thirty minutes of orientation.

Cause 6: waiting for the next model

The most intellectually respectable way to stall. There is always a more capable model arriving soon, and it is always plausible that waiting would produce a better outcome.

The flaw is that this reasoning never expires. Applied consistently it counsels permanent delay. The organisations that get value are not the ones that picked the best model; they are the ones that picked an adequate model early and spent the intervening time learning how to use it. That learning transfers when models improve. The waiting does not.

SymptomLikely causeFirst move
Pilot has quietly continued past its end dateNo exit criteriaWrite the two outcome sentences today, with a date.
Positive results but no purchaseNo budget owner involvedGet the approver into one session with the actual users.
Feedback is impressions rather than findingsTesting "AI" rather than a taskPick one workflow with a measurable before state.
Great in the pilot, ignored after rolloutUnrepresentative participantsRerun with two sceptical, fully booked colleagues.
Results judged underwhelmingNothing was configuredCorrect tier, worked examples, brief orientation, retest.
Decision deferred pending new releaseWaiting for the next modelSet the decision date first, then evaluate against it.

A pilot design that ends in a decision

Write this on one page before anything starts. If you cannot fill in a line, that gap is the thing most likely to stall you.

  • Task. One workflow, described the way the people who do it would describe it.
  • Current state. Roughly how long it takes now and what "good" looks like. A rough estimate agreed by the doers beats a precise number nobody believes.
  • Participants. Named people who do the task, including at least one sceptic.
  • Duration. Two to four weeks, with the end date in calendars from day one.
  • Setup. Which tool, which tier, which model, what worked examples they get, who gives the orientation.
  • Success looks like. One sentence, agreed by the budget owner.
  • Failure looks like. One sentence, also agreed by the budget owner.
  • Decision date and decision maker. A date and a name.

Eight lines. Nearly every stalled pilot is missing at least three of them, and it is usually the last two.

An uncomfortable but useful reframe: a pilot is not a way to reduce risk. It is a way to gather specific evidence for a specific decision. If no decision is scheduled, the pilot is not reducing risk at all, it is deferring it while spending money.

Frequently asked questions

Why do AI pilots never reach production?

Usually no exit criteria, no budget owner participating, or no defined next step. A pilot that cannot fail cannot succeed either, so it drifts until attention moves elsewhere.

How long should a pilot run?

Two to four weeks for a specific task. Longer pilots rarely produce better evidence, and they do produce more opinions, which makes the decision harder rather than easier.

What are good exit criteria?

Written before the start, naming the task, the participants, the duration, what success and failure each look like, the date of the decision, and the person who makes it.

Our pilot has already stalled. Can we restart it?

Yes, and it is usually quick, because you have already learned things. Rescope to one task, involve the budget holder, set a decision date within four weeks, and treat the previous phase as background rather than something to relitigate.

Skip the pilot purgatory

Both our packages exist to force a decision. Clarity ends with a written recommendation you can act on. Implementation ends with a configured, capped, adopted setup, live in 30 days or less, with a fixed price agreed before work starts.

Compare the packages

← Back to blog