Skip to main content
JustSoftLabJustSoftLab
JustSoftLabJustSoftLab
AI Assistant
Services/AI & GenAI/Reliability Assessment

Know how accurate your AI is, field by field.

Three weeks, fixed scope. At the end you know how accurate your AI system is today, what a wrong answer costs you, and whether swapping models is a config change or a gamble. The golden set and the eval harness stay in your repository.

3 wks

From kickoff to a measured baseline

Per field

Accuracy per field, not one blended score

Yours

Golden set and harness live in your repo

Fixed

Scoped and priced before we start

What we build

What we measure, and what you keep.

Golden set

100 to 300 examples drawn from your real production traffic, labeled field by field with the correct output. This is the asset most teams never build, and the one everything else depends on.

Eval harness

Automated runs in your stack, wired into CI. One command tells you whether a prompt change, a model swap, or a new retrieval strategy made things better or worse.

Failure taxonomy

Invented field, dropped requirement, wrong format, silent truncation, refusal. Each class counted and weighted by what it costs you when it reaches a customer.

Model comparison

Two or three candidate models run against the same golden set. Accuracy against cost against latency, in one table, so the next upgrade is a decision instead of a leap of faith.

Guardrails and fallbacks

Schema validation, behaviour on invalid output, retry and fallback policy, and the confidence threshold where the system hands the case to a human.

Ship or no-ship verdict

A written recommendation with the specific gaps behind it. If the system is ready for regulated users, you can show why. If it is not, you know exactly what to fix first.

Sound familiar?

The conversations that start this work.

A better model shipped last month and we still cannot upgrade.

You have no regression baseline, so every swap is a gamble. Once a golden set exists, changing models becomes a config change and a green test run.

Legal asked how we know the output is correct. We had no answer.

Field-level accuracy on a labeled set is the answer. It turns a claim about quality into a number your compliance team can put in a document.

It scored well in testing and users still complain.

Aggregate scores hide the failures that matter. We measure per field, so a system that is 94% correct overall but wrong on the one field that drives a decision stops looking healthy.

Common questions

Reliability Assessment, answered.

What is an AI evaluation harness?

An evaluation harness is automated test infrastructure for a system whose output is not deterministic. It runs your AI system against a fixed set of labeled examples, scores each output field by field, and reports where accuracy moved. It is the difference between believing a change helped and knowing it did.

How is this different from the testing we already do?

Conventional tests assert exact outputs, which an LLM will never reproduce twice. Evaluation measures how close the output is to a known-correct answer across many examples, and tracks that number over time. Both matter, and they answer different questions.

We already collect user feedback. Do we still need a golden set?

User feedback tells you when something went badly enough that a person complained. It misses the quiet failures: a plausible value in a field nobody checked, a dropped requirement, a subtly wrong figure. A golden set catches those before a customer sees them.

How long does the assessment take?

Three weeks for one system covering up to two task types. A one-week triage version covers a single task and produces the report without the guardrail work. Larger scopes, such as several systems or an evidence pack for a regulator, run four to five weeks.

What do we own at the end?

Everything. The golden set, the eval harness, the guardrail configuration and the report live in your repository under your license. The point of the engagement is that your team can run it after we leave, without us.

Tech stack

Tools we use in production.

Python
TypeScript
pytest
Pydantic
JSON Schema
LangGraph
OpenAI
Claude
Hugging Face
PostgreSQL
pgvector
Redis
GitHub Actions
OpenTelemetry
Grafana

Ready to build

Find out what your AI actually gets wrong.

45 minutes with our engineers. We look at your system, agree which decision points carry real cost, and scope the assessment at a fixed price. If your setup does not need this yet, we will say so on the call.