The demo went well. The model summarized a stack of contracts, answered questions about the product catalog and drafted a few support replies that read better than the real ones. Leadership liked it. A budget appeared.
Six months later it is still a demo. Nobody killed it. It just never shipped.
If you have spent any time around AI projects, you have seen this happen. The reasons are rarely about the model. They are about scope, data, ownership and plumbing, and almost all of them can be dealt with before anyone writes a line of code.
A demo answers a different question
A demo proves that a model can do a task once, on inputs someone picked, while the person who built it watches. Production asks a much harder question: can it do the task thousands of times, on inputs nobody picked, inside systems that already exist, for people who never asked for it?
Pilots are usually built to answer the first question, and teams are then surprised when the answer does not carry over. The projects that ship are built to answer the second one from the start. Everything below follows from that.
Scope it to one workflow
Stalled pilots tend to start with a vision. “An assistant for the whole company.” “Automate operations.” Visions are fine in a strategy deck. In a build plan they are a trap, because there is no way to tell when you are done.
Pilots that ship start with one workflow. Not a department, not a platform: one repeatable piece of work with a clear input, a clear output and a person who does it by hand today. Triaging inbound support tickets. Pulling fields out of supplier invoices. Drafting answers to security questionnaires from past answers.
A useful test: can you describe the workflow in two sentences, and can you name the person who will notice if it breaks? If not, the scope is still too wide.
Narrow scope is not a lack of ambition. It is how you get something real in front of users fast enough to learn from it. The second workflow also goes much faster than the first, because the integration, logging and evaluation work carries over.
Get real data in during week one
Demo data is clean. Real data is not. Real support tickets have been forwarded three times, are half in another language, and bury the actual question under a signature block. Real invoices are scanned at an angle. Real wikis contradict themselves, and nobody is sure which page is current.
A pilot built on hand-picked samples looks great until it meets the real distribution, and by then the architecture is already set. We push to get a representative slice of real data, with proper access controls, into the build in the first week. If security review is going to take a month, that is worth knowing on day one, not day sixty.
Real data also kills bad ideas early. Sometimes you discover that the information the model needs is not written down anywhere, or lives in one person’s head. Better to learn that before the build than after the launch.
Write down what “good” means before you build
Ask five stakeholders whether a pilot works and you will get five answers, each based on the last example they happened to see. That is not a basis for a go-live decision.
Projects that ship have an evaluation set: a few dozen to a few hundred real examples with known good outputs, agreed with the people who do the work today. Every change to the prompt, the model, the retrieval or the tools runs against it. You get a score that moves and a list of specific failures to look at.
It does not need to be elaborate. A spreadsheet of inputs and expected outputs plus a script that scores them beats any number of meetings. Some checks are exact, like whether it picked the right category. Some need a rubric. Some need a person to read the output. That is fine. What matters is that the bar is written down before anyone starts arguing about whether it has been cleared.
If nobody can say what good output looks like, nobody can say when the pilot is ready. So it never is.
Give it two owners
Pilots often live in an innovation team, or with whoever was most excited about AI that quarter. When it is time to go live, nobody in the business owns the outcome and nobody in engineering owns the system. So it sits.
Every project that ships has two named owners. A business owner who is accountable for the workflow and decides what good enough means. A technical owner who gets the call when it breaks. If you cannot fill both seats before the build starts, fill them before you do anything else.
Build it into the tools people already use
A lot of pilots are standalone web apps: a new tab, a new login, a chat box. People try it for a week, then drift back to the tools they spend their day in.
Adoption follows the existing workflow. If the support team works in a helpdesk, the AI draft should appear in the helpdesk. If finance lives in the ERP, the extracted invoice fields should land in the ERP, flagged for review. That means authentication, permissions and writeback are part of the core build, not a phase two.
This is usually where most of the effort goes. Calling a model takes a few lines of code. Reading from and writing back to an old system of record, safely and with the right permissions, is most of the project. Plan for it, and staff for it.
Set cost and latency budgets up front
A pilot that costs a few cents per run and takes twenty seconds feels fine when ten people are trying it. At production volume, both numbers start to matter.
Write the budgets down at the start. For example: a ticket triage step should cost a small fraction of what a minute of a support rep’s time costs, and it should finish before anyone opens the ticket. A nightly document job can be slow but has a hard monthly ceiling. Constraints like these drive real design decisions: which model handles which step, what gets cached, what gets batched, and when a smaller model with better retrieval beats a bigger model on its own.
Without budgets, teams tend to default to the largest model for everything, then discover the bill or the wait time at launch and have to rework the design.
A checklist before you start
- One workflow, described in two sentences, with a clear input and output.
- A representative sample of real data, with access approved.
- An evaluation set agreed with the people who do the work today.
- A named business owner and a named technical owner.
- A clear answer to where the output lands in existing tools.
- A cost per run and a latency target, written down.
- A way to switch it off and fall back to the manual process.
None of this is exotic. It is ordinary engineering discipline, applied earlier than teams usually apply it. The model is rarely what decides whether a pilot becomes a product. The system around it does.
If you have a pilot that has stalled, or a workflow you want to get right the first time, we’re happy to talk it through. Book a discovery call with Kryloq whenever it suits you.