BLOG · GUIDE ·

LLM evaluation for a company platform: regression testing answers before you change the model

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A company LLM platform changes its model, precision, system prompt, chunking, embedding model and serving engine several times a year, and each change should pass the same regression set and written release gates before users see it
  • Ragas defines faithfulness as the share of claims in an answer that the retrieved context supports, and context recall as the share of claims in the reference answer that the retrieved context supports; response relevancy checks alignment with the question “without evaluating factual accuracy”
  • Extraction and format cases need no judge: exact match returns 1 or 0, and promptfoo’s equals, contains, is-json and regex assertions run without a model
  • In the NeurIPS 2023 paper “Judging LLM-as-a-Judge”, strong judges such as GPT-4 reach over 80 per cent agreement with human preferences, the level found between humans, and show position and verbosity bias, so the judge is pinned, run with swapped order and checked against a human review sample
  • After the offline gates, shadow traffic sends a copy of live requests to the candidate and discards its responses, and a canary release in KServe’s serverless mode routes a set percentage of traffic to the new revision and can roll back to the previous one

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

LLM evaluation before a model change: what a platform needs

LLM evaluation on a company platform means running a fixed regression set against every change before it reaches users, and releasing the change only when it passes gates written down in advance. A platform that serves 500 to 2,000 staff changes several times a year: a new model version, a switch from FP8 to 4-bit weights, a serving engine upgrade, a new system prompt, a different chunking rule or embedding model. Each can change answers that were right last month.

The regression set is the evaluation set from the pilot, kept and extended. How to build it from users’ questions, with accepted answers and source passages, is covered in our guide to RAG on company data, and the pilot’s quality and load measurements in what to measure in a private AI pilot. This article covers the tests after go-live, from metrics and judges to a staged release.

Which platform changes need which tests

A new embedding model changes retrieval and needs a full reindex, while a new system prompt leaves retrieval alone, so not every change needs the full procedure.

CHANGEWHAT CAN BREAKWHAT TO RUN
New model or versionfacts, tone, refusals, output formatfull set, judge metrics, human sample, shadow, canary
Lower weight precisionhard cases, numbers, long answersfull set; benchmark tasks before and after
Serving engine upgradedefaults, stop conditions, output handlingfull set, deterministic checks, finish reasons
System promptrefusals, citation style, formatfull set, deterministic checks, human sample
Chunking or embedding modelwhich passages are retrievedcontext precision and recall after reindex
Judge or evaluator versionscores move with no platform changebaseline on production first

Our planning example; metric names from the Ragas and promptfoo documentation, read on 10 October 2026.

The last row concerns the evaluation tooling itself. A score is a property of the platform and of the judge together, so a new judge model or a new release of the evaluation library resets the baseline. Ragas’ documentation marks its older metrics API for deprecation in version 0.4 and removal in 1.0, which is reason enough to pin the evaluator version in the test environment.

RAG evaluation metrics: faithfulness, relevancy, precision, recall

Ragas, an open-source evaluation library under the Apache 2.0 licence, defines four metrics for the answer and the retrieval of a RAG pipeline. Faithfulness “measures how factually consistent a response is with the retrieved context”: an LLM splits the answer into claims and checks each against the retrieved passages, and the score is the share of claims supported. Response relevancy (answer relevancy in the page’s heading) generates three questions from the answer by default and averages their cosine similarity to the user’s question. Its documentation says it works “without evaluating factual accuracy”, so a fluent wrong answer can score well on it.

Context precision checks “the retriever’s ability to rank relevant chunks higher than irrelevant ones” and needs a reference answer; its variant ContextUtilization compares the passages with the generated response instead. Context recall needs a reference in every variant: in the LLM-based one the reference answer is split into claims, and the score is the share that the retrieved context supports; other variants compare reference passages or document IDs. Factual correctness compares the answer’s claims with the reference’s and reports precision, recall or F1, with F1 as the default.

METRICWHAT IT CATCHESREFERENCE NEEDEDTOOL
Faithfulnessclaims the retrieved passages do not supportnoRagas; promptfoo context-faithfulness
Response relevancyanswers that drift from the question or pad itnoRagas; promptfoo answer-relevance
Context precisionrelevant passages ranked below irrelevant onesyes, or responseRagas
Context recallreference facts the retrieval missedyesRagas; promptfoo context-recall
Factual correctnessclaims that contradict the accepted answeryesRagas; promptfoo factuality
Exact matchchanged fields, codes and numbers in extractionyesRagas; promptfoo equals
Format checksbroken JSON, missing citations, wrong structurenopromptfoo is-json, regex, contains
Benchmark tasksgeneral capability lost after quantisationbuilt inlm-evaluation-harness

Ragas metric pages, promptfoo assertion reference and the lm-evaluation-harness README, read on 10 October 2026.

Read together, faithfulness and context recall show where a failure starts. Low recall means the retrieval did not bring back what the answer needed, and the fix lies in chunking, metadata or the embedding model. High recall with low faithfulness means the model added claims of its own, which points to the model, the prompt or the precision.

Exact match and deterministic checks for extraction and format

Many cases on a company platform have one right answer: an invoice number, a contract date, a cost centre, a JSON object for the next system. For these, the Ragas Exact Match metric “returns 1 if the response is an exact match with the reference, and 0 otherwise”, and String Presence returns 1 when the answer contains the reference. Promptfoo has the same checks as assertions (equals, contains, regex, is-json), and every assertion can be negated with a not- prefix, so not-contains catches a forbidden phrase.

Deterministic checks need no judge and fail the release on a single broken case. We suggest a must-pass subset built from them: extraction fields, valid JSON, a citation on every RAG answer, and refusals for questions the asking user has no right to answer. A model change that breaks one of these cases does not go further, whatever its average scores show.

LLM-as-a-judge and its published biases

Faithfulness, relevancy and rubric scores need a model that reads the answer and grades it. The paper “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, presented at NeurIPS 2023, reports that strong judges such as GPT-4 reach “over 80% agreement” with human preferences, “the same level of agreement between humans”. That figure counts only votes without a tie; with ties counted, the paper’s Table 5 gives 66 per cent for GPT-4 against expert votes on MT-Bench and 63 to 67 per cent between experts. The paper names position, verbosity and self-enhancement biases and limited reasoning ability. Verbosity bias is defined as a judge favouring “longer, verbose responses, even if they are not as clear”. For self-enhancement, judges favouring “the answers generated by themselves”, the authors write that limited data did not let them determine whether the models show it.

In pairwise comparisons, call the judge twice with the order of the two answers swapped and count a win only if both calls agree. For questions with a known answer, use a reference-guided judge, which sees a reference answer in its prompt; in a regression set that is the accepted answer. As a precaution against self-enhancement, choose a judge from a different model family than the candidate where possible.

On a private platform the judge sees the same documents as the model it grades, so it runs on-premise too. Promptfoo’s documentation says that “by default, model-graded asserts use promptfoo’s built-in grading provider”, chosen from the credentials in the environment. Set the judge explicitly with --grader or defaultTest.options.provider; its documentation names self-hosted OpenAI-compatible judges such as vLLM. Ragas needs a judge model for its LLM-based metrics as well, and the judge’s version is pinned like the evaluator’s.

Human review sample and NIST AI 600-1

The judge’s scores are checked against people. For each release, domain experts review a sample of answers, chosen from the cases where the old and new versions differ most, without knowing which version wrote which answer. Where reviewers and the judge disagree on many cases, the judge prompt or the judge model needs work before its scores gate anything.

NIST’s Generative AI Profile (NIST AI 600-1, July 2024) describes this combination. Its suggested action MP-2.3-001 begins “Assess the accuracy, quality, reliability, and authenticity of GAI output” and lists comparison with known ground truth, human oversight and automated evaluation among the methods. MS-2.5-003, “Review and verify sources and citations in GAI system outputs”, applies before deployment and during ongoing monitoring. For a RAG assistant this means reviewers open the cited passage, not only the answer.

Our Private AI/ML service starts with a pilot on one process with clear metrics and measures the result at checkpoints. Tell us which model changes your platform has planned and how answers are checked today.

Tools: Ragas, promptfoo and lm-evaluation-harness

The three tools overlap in part. Ragas scores RAG pipelines and agents from records of question, retrieved passages, answer and reference, with LLM-based metrics and traditional ones such as exact match, BLEU and ROUGE. Promptfoo runs the prompts, providers and test cases defined in its configuration file, promptfooconfig.yaml, and combines deterministic and model-graded assertions, with a weight per assertion and an optional pass threshold per test case. Its CI/CD documentation shows quality gates that fail a pipeline when tests fail or when the pass rate falls below a set value, which turns the regression set into a step of the release pipeline.

EleutherAI’s lm-evaluation-harness, under the MIT licence, describes itself as “a framework for few-shot evaluation of language models” with over 60 standard academic benchmarks. It runs against vLLM directly or against an OpenAI-compatible server with the model types local-completions and local-chat-completions, and --log_samples keeps every response for later analysis. The harness shows whether a quantised or upgraded model kept its general capability on public tasks; public benchmarks do not contain your company’s questions, so the regression set still gates the release. Whether a new open model’s licence allows your use is a separate check, covered in open LLM licences for company use.

Release gates for a model change: an example

Gates are written before the candidate runs, so the result cannot shape them. Measure the noise first: run the regression set twice against production, because generation is not fully reproducible, and the spread between the two runs is the smallest difference a gate can detect. Our example for a platform with a set of 300 cases, of which 60 form the must-pass subset, with figures to adapt:

  1. Must-pass subset: no case that passes in production may fail.
  2. Faithfulness and factual correctness: the mean may not fall below the production baseline by more than the measured run-to-run spread.
  3. Context precision and recall: required only when retrieval changed, against the same baseline rule.
  4. Case review: every case whose score dropped is listed and read.
  5. Human sample: reviewers prefer the new answer or rate both equal in at least as many cases as they prefer the old one.
  6. Shadow traffic: the candidate answers a share of live requests with no rise in answers cut off at the length limit and no new error type.
  7. Canary: a small share of users gets the new model, with an agreed rollback trigger in the monitoring.

Istio’s documentation describes mirroring as sending “a copy of live traffic to a mirrored service”, out of band of the primary request path, with the mirrored responses discarded. To score them, the shadow service logs its answers under the same access rules as the query log, and its retrieval applies the same permission filter as production. Metrics that need no reference answer, such as faithfulness and response relevancy, work on this traffic, since live questions have no accepted answer. Requests that trigger actions, such as agent tool calls, stay out of the mirror or reach only test systems, because the shadow service would run them a second time.

KServe’s canary rollout allows “for a new version of an InferenceService to receive a percentage of traffic”, set with canaryTrafficPercent, and can be set to route all traffic back to the previous revision when a rollout step fails. Its documentation says the canary strategy “is only supported in serverless deployment mode”; in raw deployment mode the split needs weighted routes in a service mesh or gateway. During the canary, LLM monitoring with vLLM metrics shows queue, latency and finish reasons for both versions.

Our engineering partner Vixen.UNO builds the platform on Kubernetes with KServe and Kubeflow, with logging of queries and answers. Describe your serving stack and the next change in the form below.

Where failures go after a release

Every failure that users report after a release becomes a new case in the regression set, with its accepted answer and source. When answers fail on style, format or task behaviour rather than facts, a fine-tune may be the next step, and our comparison of fine-tuning and RAG shows which problem each one fixes. A fine-tuned model goes through the same gates as any other model change.

What we do

Private AI/ML is our service for private LLMs, RAG assistants and MLOps on Kubernetes, with engineering by our partner Vixen.UNO. It starts with a pilot on one process with clear metrics, measures the result at checkpoints and scales only what has proved its value. The platform logs queries and answers, so that security and legal see who accesses what, and your team is trained to run and develop it. Changes run in agreed maintenance windows with a rollback plan, with support under an agreed SLA.

FAQ

How do you evaluate an LLM on company data before changing the model?
Run a fixed regression set of company questions with accepted answers against the current model and the candidate, and compare the results case by case. Score deterministic cases with exact match and format checks, RAG answers with metrics such as faithfulness and context recall, and review a sample by hand. Release the candidate only when it passes gates written down before the test, then route traffic to it in stages.
What is LLM regression testing?
LLM regression testing reruns the same set of test cases after every change to a model, prompt, retrieval pipeline or serving engine and flags the cases that got worse. Because generation is not fully reproducible, the set is run twice against production first to measure the normal spread between runs. A drop larger than that spread, or any failure in a must-pass subset, stops the release.
What are the main RAG evaluation metrics?
Ragas defines faithfulness as the share of claims in an answer that the retrieved context supports, and response relevancy as how well the answer aligns with the question, without checking facts. Context precision measures whether relevant passages are ranked above irrelevant ones, and context recall the share of claims in the reference answer that the retrieved context supports. Factual correctness compares the answer’s claims with a reference answer.
What is Ragas used for?
Ragas is an open-source library for evaluating RAG pipelines and agents. It scores records of question, retrieved passages, answer and reference with LLM-based metrics such as faithfulness and context recall and with traditional metrics such as exact match, BLEU and ROUGE. Its LLM-based metrics need a judge model, which on a private platform runs on-premise.
Is LLM-as-a-judge reliable?
The NeurIPS 2023 paper “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena” found that strong judges such as GPT-4 agree with human preferences in over 80 per cent of votes without a tie, the same level as humans agree with each other. It documented position and verbosity bias and could not determine whether self-enhancement bias exists. Swapping answer order, giving the judge a reference answer, pinning the judge version and checking it against a human review sample reduce these effects.
Which tools are used for LLM evaluation?
Ragas scores RAG pipelines with metrics such as faithfulness and context precision. Promptfoo runs test cases from a configuration file with deterministic and model-graded assertions and can fail a CI/CD pipeline when tests fail. EleutherAI’s lm-evaluation-harness implements over 60 public benchmarks, runs them against a model or an OpenAI-compatible endpoint and shows whether a quantised or upgraded model kept its general capability.

Send us the models and serving engine your platform runs, the changes you plan for the coming months and how answers are checked today. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna