Skip to content

Guides

Evaluating an AI agent proof of concept before you scale it

· 7 min read

Judge an AI agent proof of concept against success criteria written before it was built, on a fixed set of real cases that includes the hard ones. Measure accuracy, cost per task and latency, check what data and tools it can reach and how it fails, and scale only when the results hold up outside the demo.

A convincing demo is not evidence

A demo shows an AI agent at its best, because it runs on examples the builders chose and rehearsed. The question before scaling is different: how often does the agent get real work right, what does each task cost, and what happens on the days it gets things wrong?

Treat the proof of concept as an experiment that ends in a decision: scale, change or stop. That decision needs criteria written before the results come in; otherwise almost any result can be read as promising.

Write the success criteria before the agent is built

Define what good enough means for the business, then turn it into numbers you can measure. For an agent that sorts support requests, that could be the share of requests routed to the right team, the share it correctly declines to handle and the time a person spends checking each one.

  • Task success rate on realistic cases, with a written definition of a correct result
  • The errors you can tolerate and the ones you cannot, such as a wrong answer sent to a customer
  • Cost per completed task, including model usage and the time of the person who checks the work
  • Response time that fits the workflow, depending on whether someone waits for the answer
  • The baseline: how long the task takes people today, and how often they get it right

Test on a fixed evaluation set built from real cases

Collect real inputs from your own history, record the correct outcome for each, and keep the set fixed. Include easy cases, ambiguous ones, rare ones and a few the agent should refuse. A smaller set of well-chosen cases tells you more than a large set of easy ones.

Run the agent on the whole set every time the prompt, the model, the tools or the data change, and compare the results with the previous run. Model outputs vary between runs, so run each case more than once and look at consistency, not only the best answer. Keep the cases out of the prompt and out of any examples the agent sees, or the score will flatter it.

Check how you score, too. Automatic checks work for structured outputs. Free text usually needs a person, or a second model whose judgments you have compared with a person's on a sample.

Measure cost and latency at the volume you expect

A proof of concept that handles a few dozen requests a day can hide costs that matter at thousands. Record the model calls, tokens and tool calls per task and the time each task takes. Agents that loop, retry or read long documents can cost far more than average on some inputs, so study the slowest and most expensive cases, not only the mean.

Then project the cost at your expected volume and compare it with what the work costs today, including the time people will still spend reviewing. Set hard limits per task on steps, tokens and time so that one bad input cannot run up a large bill. The OWASP Top 10 for LLM Applications lists this risk as unbounded consumption.

Design human review deliberately, then measure it

Human review is part of the design, not a temporary safety net. Decide which actions the agent may take on its own, which it may only propose and which it must never take. Anything irreversible or visible to customers, such as sending a message, changing a record or spending money, should wait for a person's approval until the agent has a long track record.

Measure the review itself: how long it takes, how often reviewers change the output and how often they approve without really checking. If reviewing takes nearly as long as doing the task, the agent is not saving time yet. If your use case could count as high-risk under the EU AI Act, effective human oversight is a legal requirement under Article 14, not a design preference, so involve your legal advisers early.

Set data and security boundaries before you scale

Scaling brings more data, more users and more tools, which is when weak boundaries start to matter. The OWASP Top 10 for LLM Applications ranks prompt injection first: text the agent reads, such as an email, a ticket or a web page, can carry instructions that redirect it. It also lists excessive agency, meaning more functions, permissions or autonomy than the task needs. Settle these points before the pilot grows.

  • Which data the agent can read, and whether personal or confidential data goes to a model provider, under which contract and retention terms
  • Which tools it can call, with the narrowest permissions and separate credentials for each agent
  • What it can change, and whether every change can be traced and reversed
  • What is logged: every tool call with its inputs and outputs, without leaking personal data
  • How each customer's or team's data is kept apart from the others'

Study how it fails, then make the call

Before deciding, read the failures, not only the score. Sort them by type: confident wrong answers, correct refusals, missed steps, wrong tool use, timeouts. An agent that fails by asking for help is far easier to deploy than one that fails silently with a plausible answer.

Then decide. Scale if the criteria are met on the evaluation set and the remaining failures are ones your process can absorb. Narrow the scope if the agent does well on only part of the task. Stop if it does not beat the baseline, and keep the evaluation set for the next attempt. Our AI agent projects follow the same sequence: one task, real examples, review by your team, and more scope only once the agent has earned trust. If you have a proof of concept to assess, you can describe it through our Start a project form.

Key takeaways

  • Write success criteria and a baseline before the agent is built, so the result can be judged honestly.
  • Test on a fixed set of real cases, including hard and ambiguous ones, and rerun it after every change.
  • Measure cost per task and latency at the expected volume, looking at the worst cases as well as the average.
  • Design human review deliberately and measure how long it takes and what it catches.
  • Set data, tool and logging boundaries before scaling, because prompt injection and excessive agency are known risks.

FAQ

How many cases does an AI agent evaluation set need?

There is no fixed number. It needs enough cases to cover the main variations of the task, the hard and ambiguous ones, and the ones the agent should refuse. Start with what you can label carefully, and add every real failure you find later.

Can another model score the agent's output?

Yes, for free-text outputs that are hard to check automatically, but only after you have compared its judgments with a person's on a sample of cases. Check that agreement again whenever you change the model or the prompt.

When should we stop an AI agent proof of concept?

When it does not beat the current way of doing the task on your success criteria, or when the review it needs costs as much time as it saves. A clear stop is a useful result: keep the evaluation set, because a newer model or a narrower task may pass it later.

Tell us what you need.

Something to build, people to find or a question to answer. In a 30-minute call we listen and tell you honestly how we can help, and what it would take.

Book a call

30 minutes, in French or English. Free.

Prefer writing? Send a short brief instead.