Picture a bakery in Dartmouth that tries an AI tool for answering customer emails. For two weeks it is wonderful. The owner shows it to her sister, her accountant and two other business owners at a networking breakfast. Then a wedding-cake season starts, nobody has time to check the drafts, one of them quotes a price from last year, and the tool goes quiet. Nobody decides to stop using it. It just stops.
That is the usual way an AI pilot ends. Not with a failure anyone can point to, but with a slow fade back to how things were done before. The demo worked. The daily work never changed.
What the numbers say, and what they don't
Most Canadian businesses are not deep into this yet. Statistics Canada reports that 12.2% of Canadian firms used AI to produce goods or deliver services in 2025, double the share the year before (Statistics Canada, 2026). So most firms trying AI right now are on their first or second attempt, with nobody down the hall who has done it before.
On what happens to those attempts, the most-quoted figure comes from MIT's NANDA project: in its 2025 report, 95% of the 300 generative AI projects it analyzed showed no measurable impact on profit and loss. That is a U.S.-led study of mostly large enterprises, and the sample was assembled from interviews and public cases rather than drawn at random (MIT NANDA, The GenAI Divide, 2025). Read it with that in mind. It says something narrower than "95% of AI fails": most of those projects had nobody measuring a before and after, so there was nothing to count. We find that more useful, because it is the one part a small team can fix.
Where the pilot actually dies
In our experience, it is rarely the model. The tool does what the demo showed. What goes missing is everything around it.
- Nobody owns it. A pilot started by one curious person belongs to that person. When they go on vacation, or the busy season starts, it has no second owner.
- No baseline. If you did not write down how long the task took last month, you cannot say in month three whether the pilot saved anything. Then the decision to continue is a mood.
- It runs on clean examples. Real work has the customer who writes in all capitals, the invoice in a format nobody warned you about, the exception that happens on Fridays.
- No checkpoint where a person decides. The first wrong price quoted to a customer ends the trust, and with it the pilot. A person signing off wherever money or reputation is involved is what lets everyone keep using it after the first mistake.
- No plan for the day after. Who changes the process, who trains the new hire, who switches it off if it misbehaves.
A pilot is a question, not a smaller version of the product
Here is the opinion we will defend: most pilots are set up to show that the tool is impressive, when they should be set up to find out whether it fits. Those are different tests. An impressive demo needs a good afternoon. Fit needs the tool run on real work, by the people who will use it, for long enough to see a bad week.
In the nine stages we work through, this is the gap between the roadmap and the build. Stage 7, Design & MVP Build, includes acceptance testing by the people who will actually use the thing, human checkpoints on anything the system decides, and a launch checklist. Stage 8 then picks it up with training and a plain-language runbook, including the off switch. The pilot does not skip those. It is the first small run of them.
Five questions to settle before the pilot starts
- What number will move, and what is it today? Hours per week on one task, days to answer a lead, errors per month. One number, written down before day one.
- Who owns it, and who covers for them? Two names.
- What is the real work it will be tested on? Pull twenty actual cases from the last quarter, including the ugly ones.
- Where does a person look before it goes out? Name the step.
- What happens at the end of the pilot? Decide in advance what result means expand, what means fix, and what means stop. A pilot with no end date and no decision rule drifts.
If the task is a routine one such as drafting follow-ups or keeping a pipeline tidy, the shortcut is to skip the build and start from a ready-made setup, which we wrote about in what Claude's Small Business plugin actually does. Pilots that need your own data, your own systems or a decision the software makes on your behalf are a different job, and that is the build work this post is about. The other half of this, trusting the machine too much, is in AI needs a pilot.
If you have a pilot running now and are not sure which of the five questions it would fail, a free 30-minute call is enough time to go through it with you.
