Know how accurate your AI is, field by field.
Three weeks, fixed scope. At the end you know how accurate your AI system is today, what a wrong answer costs you, and whether swapping models is a config change or a gamble. The golden set and the eval harness stay in your repository.
From kickoff to a measured baseline
Accuracy per field, not one blended score
Golden set and harness live in your repo
Scoped and priced before we start
What we build
What we measure, and what you keep.
Golden set
100 to 300 examples drawn from your real production traffic, labeled field by field with the correct output. This is the asset most teams never build, and the one everything else depends on.
Eval harness
Automated runs in your stack, wired into CI. One command tells you whether a prompt change, a model swap, or a new retrieval strategy made things better or worse.
Failure taxonomy
Invented field, dropped requirement, wrong format, silent truncation, refusal. Each class counted and weighted by what it costs you when it reaches a customer.
Model comparison
Two or three candidate models run against the same golden set. Accuracy against cost against latency, in one table, so the next upgrade is a decision instead of a leap of faith.
Guardrails and fallbacks
Schema validation, behaviour on invalid output, retry and fallback policy, and the confidence threshold where the system hands the case to a human.
Ship or no-ship verdict
A written recommendation with the specific gaps behind it. If the system is ready for regulated users, you can show why. If it is not, you know exactly what to fix first.
Sound familiar?
The conversations that start this work.
“A better model shipped last month and we still cannot upgrade.”
You have no regression baseline, so every swap is a gamble. Once a golden set exists, changing models becomes a config change and a green test run.
“Legal asked how we know the output is correct. We had no answer.”
Field-level accuracy on a labeled set is the answer. It turns a claim about quality into a number your compliance team can put in a document.
“It scored well in testing and users still complain.”
Aggregate scores hide the failures that matter. We measure per field, so a system that is 94% correct overall but wrong on the one field that drives a decision stops looking healthy.
Common questions
Reliability Assessment, answered.
What is an AI evaluation harness?
An evaluation harness is automated test infrastructure for a system whose output is not deterministic. It runs your AI system against a fixed set of labeled examples, scores each output field by field, and reports where accuracy moved. It is the difference between believing a change helped and knowing it did.
How is this different from the testing we already do?
Conventional tests assert exact outputs, which an LLM will never reproduce twice. Evaluation measures how close the output is to a known-correct answer across many examples, and tracks that number over time. Both matter, and they answer different questions.
We already collect user feedback. Do we still need a golden set?
User feedback tells you when something went badly enough that a person complained. It misses the quiet failures: a plausible value in a field nobody checked, a dropped requirement, a subtly wrong figure. A golden set catches those before a customer sees them.
How long does the assessment take?
Three weeks for one system covering up to two task types. A one-week triage version covers a single task and produces the report without the guardrail work. Larger scopes, such as several systems or an evidence pack for a regulator, run four to five weeks.
What do we own at the end?
Everything. The golden set, the eval harness, the guardrail configuration and the report live in your repository under your license. The point of the engagement is that your team can run it after we leave, without us.
Tech stack
Tools we use in production.
Ready to build
Find out what your AI actually gets wrong.
45 minutes with our engineers. We look at your system, agree which decision points carry real cost, and scope the assessment at a fixed price. If your setup does not need this yet, we will say so on the call.
AI projects we delivered





