Kill Your Pilot
AI pilots without an end date become permanent budget lines that never produce a decision. Every pilot needs graduation criteria, a kill date, and a named owner accountable for calling it.
I was on a call last month with a VP of engineering who told me, with no apparent irony, that his team was "eighteen months into a six-month AI pilot." Nobody could tell me what the pilot was supposed to prove anymore. The vendor invoice renewed quarterly. A dozen engineers used the tool sometimes. Leadership's official position was "we're still evaluating."
That's not a pilot. That's a subscription with extra steps.
The pilot that never ends
The pattern is everywhere, and it's worth naming precisely because it feels like progress. A pilot gives everyone something they want: the CTO gets to say the org is "doing AI," the vendor gets a logo and a renewal, the engineers who like the tool get to keep it, and nobody has to make a decision that could later be wrong. The one thing a perpetual pilot never produces is the thing pilots exist to produce — a decision.
The cost isn't just the invoice. It's that pilot status quietly caps the value. Nobody builds real workflow around a tool that might disappear. Nobody invests in shared configuration, context files, or training for something "we're still evaluating." So the tool underperforms, which conveniently justifies staying in evaluation mode. I've watched teams run this loop for two years.
Meanwhile the industry has moved from experimentation to rationalization. Surveys this year show roughly two-thirds of technology leaders actively cutting their AI vendor portfolios, and the experimentation budgets of 2024–2025 are getting consolidated into fewer, bigger bets. If your pilots don't produce decisions, someone in finance will eventually produce them for you — and they'll optimize for the invoice, not the outcome.
What a real pilot looks like
A pilot is an experiment, and experiments have three properties that most AI pilots are missing.
A falsifiable claim. Not "see if the team likes it," but something you could be wrong about: "Coding agents will reduce cycle time on migration work by a third for the platform team" or "an agent triaging support tickets will cut first-response time in half without increasing escalations." If you can't state what would count as failure, you're not piloting — you're shopping.
A scheduled end date with a forced decision. Ninety days is enough for a coding agent; six months is enough for almost anything. On the end date, there are exactly three allowed outcomes: graduate (roll out with real budget, ownership, and enablement), kill (cancel the contract, write down what you learned), or extend once — with a documented reason and a new end date. "Extend indefinitely" is not on the menu. The second extension request is a kill signal wearing a costume.
A named owner who is accountable for the call. Not a committee, not "the AI working group." One person whose job includes standing up at the end date and saying graduate or kill, with the evidence. If nobody wants that job, that tells you the pilot doesn't matter enough to run.
Instrument the decision, not the demo
The reason most pilots can't end is that they weren't instrumented to answer their own question. If the claim is about cycle time, you need the baseline before the pilot starts — cycle time, review time, defect escape rate for the affected work over the previous 30–60 days. Practitioners in this space consistently converge on the same advice: track AI-assisted work separately from the rest, because blended averages hide everything. A pilot that shows "PRs went up" tells you nothing if review time ballooned downstream and the bottleneck just moved.
This is boring work, which is why it gets skipped. But it's a week of setup, and it's the difference between ending the pilot with a decision and ending it with vibes.
Graduation is a project, not a checkbox
One more failure mode: pilots that succeed and then die anyway, because "graduate" was never scoped. Graduation means budget moves from experimental to operational. It means someone owns enablement — shared configuration, onboarding, internal patterns — as a job, not a hobby. It means security and procurement review happened during the pilot, not as a surprise after it. If your pilot succeeds in week 12 and legal review starts in week 13, you've built a six-month gap between "it works" and "we use it," and momentum doesn't survive that.
The teams I've seen do this well run the workstreams in parallel: while engineers evaluate the tool, someone runs the procurement, security, and rollout-planning track so that a "graduate" decision executes in weeks.
The takeaway
Audit your pilots this week. For each one, ask three questions: What claim is it testing? When does it end? Who makes the call? Any pilot missing an answer gets one of two treatments — give it real success criteria, an end date, and an owner, or kill it now and stop pretending. A killed pilot that produces a clear "no, and here's why" is worth more than a zombie pilot that produces eighteen months of "we're still evaluating."
Pilots are supposed to be how you learn fast. Don't let them become how you avoid deciding.
Wes Goldwater
Director of Engineering at Prosigliere · writing the no-hype playbook for cloud & AI.
Keep reading
Delegation Without Atrophy
An RCT found AI-assisted developers scored 17% lower on comprehension. The skills that decay are exactly the ones verification depends on — so treat skill maintenance as a managed budget.
Estimate the Review, Not the Writing
Story points measured human drafting effort, which agents just zeroed out. Work is now bimodal: size verification cost, split agent-eligible from human-led lanes, and cap agent output at review capacity.