LLM evaluation for a company platform: regression testing answers before you change the model
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A company LLM platform changes its model, precision, system prompt, chunking, embedding model and serving engine several times a year, and each change should pass the same regression set and written release gates before users see it
- Ragas defines faithfulness as the share of claims in an answer that the retrieved context supports, and context recall as the share of claims in the reference answer that the retrieved context supports; response relevancy checks alignment with the question “without evaluating factual accuracy”
- Extraction and format cases need no judge: exact match returns 1 or 0, and promptfoo’s equals, contains, is-json and regex assertions run without a model
- In the NeurIPS 2023 paper “Judging LLM-as-a-Judge”, strong judges such as GPT-4 reach over 80 per cent agreement with human preferences, the level found between humans, and show position and verbosity bias, so the judge is pinned, run with swapped order and checked against a human review sample
- After the offline gates, shadow traffic sends a copy of live requests to the candidate and discards its responses, and a canary release in KServe’s serverless mode routes a set percentage of traffic to the new revision and can roll back to the previous one
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
LLM evaluation before a model change: what a platform needs
LLM evaluation on a company platform means running a fixed regression set against every change before it reaches users, and releasing the change only when it passes gates written down in advance. A platform that serves 500 to 2,000 staff changes several times a year: a new model version, a switch from FP8 to 4-bit weights, a serving engine upgrade, a new system prompt, a different chunking rule or embedding model. Each can change answers that were right last month.
The regression set is the evaluation set from the pilot, kept and extended. How to build it from users’ questions, with accepted answers and source passages, is covered in our guide to RAG on company data, and the pilot’s quality and load measurements in what to measure in a private AI pilot. This article covers the tests after go-live, from metrics and judges to a staged release.
Which platform changes need which tests
A new embedding model changes retrieval and needs a full reindex, while a new system prompt leaves retrieval alone, so not every change needs the full procedure.
| CHANGE | WHAT CAN BREAK | WHAT TO RUN |
|---|---|---|
| New model or version | facts, tone, refusals, output format | full set, judge metrics, human sample, shadow, canary |
| Lower weight precision | hard cases, numbers, long answers | full set; benchmark tasks before and after |
| Serving engine upgrade | defaults, stop conditions, output handling | full set, deterministic checks, finish reasons |
| System prompt | refusals, citation style, format | full set, deterministic checks, human sample |
| Chunking or embedding model | which passages are retrieved | context precision and recall after reindex |
| Judge or evaluator version | scores move with no platform change | baseline on production first |
Our planning example; metric names from the Ragas and promptfoo documentation, read on 10 October 2026.
The last row concerns the evaluation tooling itself. A score is a property of the platform and of the judge together, so a new judge model or a new release of the evaluation library resets the baseline. Ragas’ documentation marks its older metrics API for deprecation in version 0.4 and removal in 1.0, which is reason enough to pin the evaluator version in the test environment.
RAG evaluation metrics: faithfulness, relevancy, precision, recall
Ragas, an open-source evaluation library under the Apache 2.0 licence, defines four metrics for the answer and the retrieval of a RAG pipeline. Faithfulness “measures how factually consistent a response is with the retrieved context”: an LLM splits the answer into claims and checks each against the retrieved passages, and the score is the share of claims supported. Response relevancy (answer relevancy in the page’s heading) generates three questions from the answer by default and averages their cosine similarity to the user’s question. Its documentation says it works “without evaluating factual accuracy”, so a fluent wrong answer can score well on it.
Context precision checks “the retriever’s ability to rank relevant chunks higher than irrelevant ones” and needs a reference answer; its variant ContextUtilization compares the passages with the generated response instead. Context recall needs a reference in every variant: in the LLM-based one the reference answer is split into claims, and the score is the share that the retrieved context supports; other variants compare reference passages or document IDs. Factual correctness compares the answer’s claims with the reference’s and reports precision, recall or F1, with F1 as the default.
| METRIC | WHAT IT CATCHES | REFERENCE NEEDED | TOOL |
|---|---|---|---|
| Faithfulness | claims the retrieved passages do not support | no | Ragas; promptfoo context-faithfulness |
| Response relevancy | answers that drift from the question or pad it | no | Ragas; promptfoo answer-relevance |
| Context precision | relevant passages ranked below irrelevant ones | yes, or response | Ragas |
| Context recall | reference facts the retrieval missed | yes | Ragas; promptfoo context-recall |
| Factual correctness | claims that contradict the accepted answer | yes | Ragas; promptfoo factuality |
| Exact match | changed fields, codes and numbers in extraction | yes | Ragas; promptfoo equals |
| Format checks | broken JSON, missing citations, wrong structure | no | promptfoo is-json, regex, contains |
| Benchmark tasks | general capability lost after quantisation | built in | lm-evaluation-harness |
Ragas metric pages, promptfoo assertion reference and the lm-evaluation-harness README, read on 10 October 2026.
Read together, faithfulness and context recall show where a failure starts. Low recall means the retrieval did not bring back what the answer needed, and the fix lies in chunking, metadata or the embedding model. High recall with low faithfulness means the model added claims of its own, which points to the model, the prompt or the precision.
Exact match and deterministic checks for extraction and format
Many cases on a company platform have one right answer: an invoice number, a contract date, a cost centre, a JSON object for the next system. For these, the Ragas Exact Match metric “returns 1 if the response is an exact match with the reference, and 0 otherwise”, and String Presence returns 1 when the answer contains the reference. Promptfoo has the same checks as assertions (equals, contains, regex, is-json), and every assertion can be negated with a not- prefix, so not-contains catches a forbidden phrase.
Deterministic checks need no judge and fail the release on a single broken case. We suggest a must-pass subset built from them: extraction fields, valid JSON, a citation on every RAG answer, and refusals for questions the asking user has no right to answer. A model change that breaks one of these cases does not go further, whatever its average scores show.
LLM-as-a-judge and its published biases
Faithfulness, relevancy and rubric scores need a model that reads the answer and grades it. The paper “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, presented at NeurIPS 2023, reports that strong judges such as GPT-4 reach “over 80% agreement” with human preferences, “the same level of agreement between humans”. That figure counts only votes without a tie; with ties counted, the paper’s Table 5 gives 66 per cent for GPT-4 against expert votes on MT-Bench and 63 to 67 per cent between experts. The paper names position, verbosity and self-enhancement biases and limited reasoning ability. Verbosity bias is defined as a judge favouring “longer, verbose responses, even if they are not as clear”. For self-enhancement, judges favouring “the answers generated by themselves”, the authors write that limited data did not let them determine whether the models show it.
In pairwise comparisons, call the judge twice with the order of the two answers swapped and count a win only if both calls agree. For questions with a known answer, use a reference-guided judge, which sees a reference answer in its prompt; in a regression set that is the accepted answer. As a precaution against self-enhancement, choose a judge from a different model family than the candidate where possible.
On a private platform the judge sees the same documents as the model it grades, so it runs on-premise too. Promptfoo’s documentation says that “by default, model-graded asserts use promptfoo’s built-in grading provider”, chosen from the credentials in the environment. Set the judge explicitly with --grader or defaultTest.options.provider; its documentation names self-hosted OpenAI-compatible judges such as vLLM. Ragas needs a judge model for its LLM-based metrics as well, and the judge’s version is pinned like the evaluator’s.
Human review sample and NIST AI 600-1
The judge’s scores are checked against people. For each release, domain experts review a sample of answers, chosen from the cases where the old and new versions differ most, without knowing which version wrote which answer. Where reviewers and the judge disagree on many cases, the judge prompt or the judge model needs work before its scores gate anything.
NIST’s Generative AI Profile (NIST AI 600-1, July 2024) describes this combination. Its suggested action MP-2.3-001 begins “Assess the accuracy, quality, reliability, and authenticity of GAI output” and lists comparison with known ground truth, human oversight and automated evaluation among the methods. MS-2.5-003, “Review and verify sources and citations in GAI system outputs”, applies before deployment and during ongoing monitoring. For a RAG assistant this means reviewers open the cited passage, not only the answer.
Our Private AI/ML service starts with a pilot on one process with clear metrics and measures the result at checkpoints. Tell us which model changes your platform has planned and how answers are checked today.
Tools: Ragas, promptfoo and lm-evaluation-harness
The three tools overlap in part. Ragas scores RAG pipelines and agents from records of question, retrieved passages, answer and reference, with LLM-based metrics and traditional ones such as exact match, BLEU and ROUGE. Promptfoo runs the prompts, providers and test cases defined in its configuration file, promptfooconfig.yaml, and combines deterministic and model-graded assertions, with a weight per assertion and an optional pass threshold per test case. Its CI/CD documentation shows quality gates that fail a pipeline when tests fail or when the pass rate falls below a set value, which turns the regression set into a step of the release pipeline.
EleutherAI’s lm-evaluation-harness, under the MIT licence, describes itself as “a framework for few-shot evaluation of language models” with over 60 standard academic benchmarks. It runs against vLLM directly or against an OpenAI-compatible server with the model types local-completions and local-chat-completions, and --log_samples keeps every response for later analysis. The harness shows whether a quantised or upgraded model kept its general capability on public tasks; public benchmarks do not contain your company’s questions, so the regression set still gates the release. Whether a new open model’s licence allows your use is a separate check, covered in open LLM licences for company use.
Release gates for a model change: an example
Gates are written before the candidate runs, so the result cannot shape them. Measure the noise first: run the regression set twice against production, because generation is not fully reproducible, and the spread between the two runs is the smallest difference a gate can detect. Our example for a platform with a set of 300 cases, of which 60 form the must-pass subset, with figures to adapt:
- Must-pass subset: no case that passes in production may fail.
- Faithfulness and factual correctness: the mean may not fall below the production baseline by more than the measured run-to-run spread.
- Context precision and recall: required only when retrieval changed, against the same baseline rule.
- Case review: every case whose score dropped is listed and read.
- Human sample: reviewers prefer the new answer or rate both equal in at least as many cases as they prefer the old one.
- Shadow traffic: the candidate answers a share of live requests with no rise in answers cut off at the length limit and no new error type.
- Canary: a small share of users gets the new model, with an agreed rollback trigger in the monitoring.
Istio’s documentation describes mirroring as sending “a copy of live traffic to a mirrored service”, out of band of the primary request path, with the mirrored responses discarded. To score them, the shadow service logs its answers under the same access rules as the query log, and its retrieval applies the same permission filter as production. Metrics that need no reference answer, such as faithfulness and response relevancy, work on this traffic, since live questions have no accepted answer. Requests that trigger actions, such as agent tool calls, stay out of the mirror or reach only test systems, because the shadow service would run them a second time.
KServe’s canary rollout allows “for a new version of an InferenceService to receive a percentage of traffic”, set with canaryTrafficPercent, and can be set to route all traffic back to the previous revision when a rollout step fails. Its documentation says the canary strategy “is only supported in serverless deployment mode”; in raw deployment mode the split needs weighted routes in a service mesh or gateway. During the canary, LLM monitoring with vLLM metrics shows queue, latency and finish reasons for both versions.
Our engineering partner Vixen.UNO builds the platform on Kubernetes with KServe and Kubeflow, with logging of queries and answers. Describe your serving stack and the next change in the form below.
Where failures go after a release
Every failure that users report after a release becomes a new case in the regression set, with its accepted answer and source. When answers fail on style, format or task behaviour rather than facts, a fine-tune may be the next step, and our comparison of fine-tuning and RAG shows which problem each one fixes. A fine-tuned model goes through the same gates as any other model change.
What we do
Private AI/ML is our service for private LLMs, RAG assistants and MLOps on Kubernetes, with engineering by our partner Vixen.UNO. It starts with a pilot on one process with clear metrics, measures the result at checkpoints and scales only what has proved its value. The platform logs queries and answers, so that security and legal see who accesses what, and your team is trained to run and develop it. Changes run in agreed maintenance windows with a rollback plan, with support under an agreed SLA.
FAQ
How do you evaluate an LLM on company data before changing the model?
What is LLM regression testing?
What are the main RAG evaluation metrics?
What is Ragas used for?
Is LLM-as-a-judge reliable?
Which tools are used for LLM evaluation?
Send us the models and serving engine your platform runs, the changes you plan for the coming months and how answers are checked today. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day