Technologies

AI evaluation and testing

A model that returns a valid answer can still return the wrong one. Evaluation is how you find out, ideally before your users or your auditors do.

Starts With
Your own cases
Separates
Consistency from accuracy
Every Run
On a pinned version
Provider Stance
No reseller agreements
The Gap

Most AI evaluations answer a different question than the one you are about to bet on.

Vendor benchmarks, public leaderboards and demo results all measure something real. They rarely measure your inputs, your error costs or the model version you will actually run. The gap between those numbers and your operation is where most disappointing pilots come from.

01

Benchmarks measure someone else’s task

A leaderboard score tells you how a model does on a public dataset. It says little about your documents, your language mix or the edge cases your process actually produces.

02

Agreement is not accuracy

Two models agreeing with each other can both be wrong. When the reference labels come from another model, the result measures similarity, not correctness.

03

A valid format is not a right answer

Structured outputs guarantee that an answer fits the schema. They do not guarantee that the category, score or decision inside it is correct.

Measurement

What a useful evaluation measures

We measure six properties separately. A system can be strong on one and weak on the next, and each kind of failure costs you something different.

An evaluation dashboard with accuracy, calibration, latency and cost panels, one showing a drop highlighted in red
  1. 01

    Accuracy against real outcomes

    The reference is what actually happened in your process, such as the team that was right, the invoice that was approved or the defect that was confirmed. Another model’s opinion is not a reference.

  2. 02

    Consistency on repeated cases

    The same inputs, run again on the same pinned model version. How far does the judgment move? Repeatability and correctness are separate questions, so they get separate tests.

  3. 03

    Calibration of the probabilities

    If a system says 80 percent, is it right about 80 percent of the time across many cases? Any threshold you set on a probability is only as good as this.

  4. 04

    The cost of errors, not just their count

    A wrong support queue and a wrong payment approval are not the same mistake. Confidently wrong answers are weighted by what they would cost, not averaged away.

  5. 05

    Review load and fallback behaviour

    How many cases go to people, what happens on a timeout or missing result, and whether unsupported inputs are recognised rather than forced into an answer.

  6. 06

    Latency and cost at production volume

    Measured end to end, including retries and human review, rather than as the price of a single model call in a demo.

Reading a Result

Five questions before you trust a number

Most claims about AI performance can be checked in ten minutes by asking what was compared, against what, and under which conditions. These are the questions we ask of vendor figures and of our own.

01

What exactly was compared, and against which reference?

02

Was the reference an established outcome, or another model’s judgment?

03

Which model version produced the result, and is that version pinned?

04

Do the test cases look like your inputs in language, length, noise and adversarial content?

05

What does a wrong answer cost in this workflow, and how was that weighted?

An ultrawide monitor showing the same set of test cases twice, side by side, with pass, warning and fail markers that differ between the two runs
The Rule of Thumb
If a result doesn’t say what it was compared against and which model version produced it, treat it as a claim rather than evidence.

The Evaluation Set

Build the evaluation set before you choose the model.

The set of cases you score every candidate against is the most durable asset in an AI project. Models get replaced every few months, and a good evaluation set stays useful through all of them.

A monitor showing a library of evaluation cases as tagged thumbnails, a few outlined in amber as deliberately hard cases
Real cases from the process
Drawn from the work as it actually arrives, including the messy, incomplete and ambiguous inputs that never appear in a demo.
Labelled against outcomes
Each case carries the answer your process confirmed, labelled by the people who own it. The model under test never supplies its own labels.
The hard cases on purpose
Arithmetic, date comparisons, long inputs with irrelevant detail, other languages and content that tries to steer the answer. Where real examples are scarce, synthetic ones fill the gap.
Vendor-neutral by design
The same set scores a native decision model, an open model and a conventional language model, so switching providers never means redesigning how you measure.

Where This Applies

What we evaluate, and what it looks like

The same discipline applies to every kind of model. The reference and the cost of an error change from case to case; the method does not.

A reliability diagram with bars close to the diagonal and a confusion-matrix heat map with a strong diagonal

Classification, routing and scoring

Decisions with a fixed set of answers, such as which team, which category, or pass or fail. We measure accuracy per class, where the errors concentrate, and whether the stated probabilities hold up well enough to set a threshold on.

Benefits:

Accuracy per class, not just overall
Calibration checked before thresholds are set
The confusions that matter made visible
A document with extracted fields outlined in green, two in red where they disagree with the confirmed value and one in amber for review

Document extraction

Fields pulled from invoices, forms and drawings, checked value by value against what the process confirmed. That includes fields left empty or put in the wrong place.

Benefits:

Field-level accuracy against confirmed values
Silent errors separated from flagged ones
A review rate you can plan around
A generated answer next to its source document, supported sentences highlighted in green and one unsupported sentence in red

Generative answers and RAG

Answers written by a language model, checked sentence by sentence against the sources they claim to draw on. Grounding, completeness and refusals are measured separately from how fluent the answer reads.

Benefits:

Field-level accuracy against confirmed values
Silent errors separated from flagged ones
A review rate you can plan around
Tablets on a conveyor, each outlined and classified as good or defective

Vision inspection models

Detection and defect models measured on the parts, lighting and cameras of your own line, with the rare defect classes weighted by what a miss would cost rather than averaged away.

Benefits:

Tested on your line, not a benchmark
Misses weighted by what they cost
Rare classes measured on their own

Define what a costly error looks like before you measure speed.

From the decision to a repeatable test

How an evaluation engagement looks

We start from the decision the system is meant to make and end with a test you can rerun on every model update.

1

Define the decision and its error cost

What is being decided, what the allowed answers are, and what each kind of mistake costs the business.

2

Build the evaluation set

Real cases labelled against confirmed outcomes, with the hard cases deliberately included.

3

Run the candidates on pinned versions

Accuracy, consistency, calibration, latency and cost, measured side by side under the same conditions.

4

Set thresholds, then rerun on every update

Review thresholds and fallback rules set from the measured results, and the test wired into the release process.

Book a Strategy & Architecture Review

Bring us the number you are being asked to trust.

Whether it’s a vendor benchmark, a pilot result or your own team’s accuracy figure, we’ll tell you what it actually measures and what a test on your cases would need.

Book a Strategy & Architecture Review
nAIxt Technologies GmbH
Am Forst 2
82166 Gräfelfing, Germany
+49 89 54196515
info@naixt-technologies.de