In a prototype, the prompt feels like the product. You tweak a sentence, the output improves, and it feels like progress. In production, the prompt is one file among many, and usually not the one that decides whether the system can be trusted.
These are the pieces we build around every agent before we call it production-ready. None of them are glamorous. All of them are what separate an agent people rely on from one they quietly stop using after its first bad week.
Evals: a test suite for behavior
An eval set is a collection of real cases with known good outcomes. It is the closest thing an AI system has to a test suite, and it should be treated like one: run on every change to the prompt, the model, the tools or the retrieval, and wired into CI so that a regression blocks a deploy.
For agents, checking the final answer is not enough. You also want to check the trajectory. Did it call the right tools? Did it pass sensible arguments? Did it stop when it had enough, or wander for twenty steps? An agent can reach a correct answer by a route that would be expensive or unsafe at scale.
Use a mix of checks. Exact matches where the output is structured. Rubrics scored by a model grader where it is not, with a person spot-checking the grader regularly. Include the awkward cases on purpose: missing data, ambiguous requests, hostile inputs, questions the agent should decline. And every time something goes wrong in production, add it to the set. Over time, the eval set becomes the most valuable artifact in the project.
Guardrails that live in code
A line in the prompt saying “never issue a refund over $500” is a suggestion. A refund tool that rejects anything over $500 is a rule. Guardrails belong in code, where the model cannot talk its way past them.
- Least-privilege tools. Read-only by default, write access only where needed, and scoped to the permissions of the user the agent is acting for.
- Validated inputs and outputs. Tool arguments and final outputs are checked against schemas before anything acts on them.
- Hard limits. A maximum number of steps and tool calls per run, and timeouts on every external call.
- Untrusted content stays untrusted. Text from emails, web pages and documents can contain instructions, and it should never be able to trigger a privileged action without an independent check.
- Deliberate data handling. Redact what the model does not need, and be careful about what ends up in logs.
Observability: see every step
When someone reports that the agent gave a strange answer yesterday afternoon, you need to be able to find that run and see exactly what happened. Without that, debugging is guesswork.
Trace every run end to end: the input, the retrieved context, each model call, each tool call with its arguments and results, token counts, latency, cost and the final output. Tag each trace with the versions of the prompt, model and tools that produced it. Whether you use an OpenTelemetry-based setup, a tool like Langfuse or LangSmith, or a well-designed table in Postgres matters less than doing it from day one.
On top of the traces, keep a small dashboard: volume, error rate, cost per run, latency, and how often the agent escalates to a person. Alert on sudden changes. A jump in escalations or in average steps per run is often the first sign that something upstream has changed.
Human-in-the-loop, designed in from the start
Decide early which actions need a person’s approval. A sensible starting rule: anything irreversible, customer-facing or above a money threshold goes through review at first.
Then make review fast. Put the approval step inside the tool reviewers already use, show the agent’s reasoning and sources next to the draft, and make editing as easy as approving. If review is slow, people will either rubber-stamp it or route around it.
Capture every edit and rejection. They are the best signal you will get about where the agent falls short, and they turn directly into new eval cases. As evidence builds for a given type of action, you can loosen approval for that action and keep it for the rest.
The agent also needs a clean way to give up. When it is unsure, it should hand off to a person with a summary of what it tried, not guess.
An agent that knows when to stop and ask is worth more than one that is right slightly more often.
Cost controls
Agent costs vary from run to run, which means they can surprise you. Put limits in place before you need them: a token and tool-call budget per run, quotas per user or tenant, and an alert when spend moves outside its normal range.
Then design for efficiency. Route easy steps to smaller, cheaper models and save the larger ones for the steps that need them. Use prompt caching for long, stable context. Batch work that does not need an immediate answer.
Measure cost per successful outcome, not per call. For example, if a run costs a few cents and resolves a ticket that would take someone ten minutes, that is a good trade. If a bug in the loop makes the same run cost a hundred times more, you want a hard stop and an alert, not a surprise at the end of the month.
Version everything that changes behavior
Prompts, tool definitions, retrieval settings, model identifiers and the eval set all change how the agent behaves, so they all belong in version control, together. Pin exact model versions rather than pointing at whatever is newest, and re-run the evals before you move to a new one.
Roll out changes the way you would any risky deploy. Run the new version in shadow mode against live traffic, compare the results, release it to a small slice of users, and keep rollback to a single step. Because every trace records which versions produced it, you can tell exactly which change caused a shift in behavior.
A handover the client team can own
We treat an engagement as finished when the client’s team can run and change the system without us. If only the people who built it can safely touch it, it is not done.
In practice, the handover includes:
- Code in the client’s repositories and infrastructure in the client’s cloud accounts, from the first day.
- A runbook covering the common failure modes and what to do about each one.
- The eval suite, with a walkthrough of how to add cases and read the results.
- Dashboards and alerts connected to the client’s own on-call process.
- A short decision log explaining why the system is built the way it is.
- Working sessions where the client’s engineers make a real change, from eval to deploy, while we pair with them.
The prompt still matters
None of this makes the prompt unimportant. A good prompt is still the difference between an agent that understands the task and one that does not. It is just the start of the work, not the whole of it.
If you are planning an agent and want these pieces in place from the start, we’d be glad to talk. Book a discovery call with Kryloq and we’ll walk through your use case.