[Virtual event] Automate Your Software Factory - Nov 5 — Save my seat

Skip to:
BlogRight arrowBest Practices
Right arrowLLM evaluations: Tutorial and best practices

Jan 6, 2026

LLM evaluations: Tutorial and best practices

Learn how to evaluate large language models before you ship them and monitor their performance continuously in production.

LaunchDarkly
LaunchDarkly
LLM evaluation: scoring models offline against a fixed dataset before release and continuously on live production traffic

Key Takeaways

  • Offline evaluation scores a candidate against a fixed dataset before release; online evaluation scores live production responses continuously, on traffic the test set never contained.
  • LLM judges need evaluation, too: RAND research from March 2026 found that the weakest of four frontier judges matched the expected score only 40% of the time on paraphrased responses.
  • A reliable production evaluation loop runs in five stages: detect, attribute, bound, act, and verify.
  • Treating prompts, models, and parameters as versioned configuration lets a failing evaluation score roll back production behavior at runtime instead of through a deploy cycle.

Evaluating a model before deployment tells you how it performed on a test set, not how it will respond to a customer an hour after it goes live.

AI behavior is probabilistic. The same prompt can return 20 different answers, and the model behind it can be updated without notice. A score that cleared your threshold in an offline run is a statement about a fixed dataset at a specific moment. It doesn't guarantee anything about live traffic.

This is the gap most evaluation practices still leave open. Your teams build careful test sets, run judges, chart the results, and configure alerts. Then when an alert fires, a person still has to notice it, diagnose it, and determine what to do—all while the model keeps answering.

The table below summarizes the concepts this article explores in detail.

Key LLM evaluation concepts

Concept

Description

LLM evaluation

The process of measuring how well a large language model performs across tasks, contexts, and safety dimensions.

Offline evaluation

Scoring a candidate model, prompt, or configuration against a fixed dataset before it reaches production traffic.

Online evaluation

Scoring live production responses continuously, on traffic and inputs the test set did not contain.

Perplexity

A metric for intrinsic evaluation that measures how surprised a model is by the actual next word in a sequence. Lower is better.

BLEU / ROUGE / METEOR

Surface-form metrics for comparing generated text to references; used in translation, summarization, and generation tasks.

LLM-as-a-judge

Using a powerful LLM to evaluate or grade the outputs of other models or tasks.

Judge reliability

The degree to which a judge returns the same verdict on the same input. This is a limiting factor on what you can safely automate.

Red teaming

A method from security and military strategy: probing a model with adversarial inputs to find vulnerabilities or unsafe behavior.

Jailbreaking

Using clever or adversarial prompts to bypass an LLM's built-in safety or guardrails.

Null models

A research concept demonstrating that trivial or constant-response models can sometimes skew benchmarks, revealing flaws in tests.

What is LLM evaluation, and why is it challenging?

You can evaluate either an LLM itself or an LLM system.

In model evaluation, engineers assess a particular model across tasks or domains without assistive components such as retrieval or tools. System evaluation assesses a particular setup, such as a chatbot built on an open-weight model for mental health support.

The stochastic nature of generated text makes this difficult. Three different people would get different responses to the prompt, "Explain quantum mechanics to me as an 8-year-old," and all three responses could be accurate.

Unfortunately, metrics designed for tasks with a single correct output, like the correct label for an image, are not as useful for evaluating open-ended language generation. LLMs can have more than one "right" answer, so LLM evaluation requires judging things like correctness, relevance, and instruction following instead of straightforward output matching.

There is also a philosophical problem underneath the technical one, because correctness can sometimes be subjective. For example, the right answer to a prompt like "What is the right time for dinner?" depends more on the audience and context than on the model.

Coherence, helpfulness, prompt sensitivity, and safety all have to be balanced, and each demands its own measure. The shortage of good benchmarks and the ease with which existing ones can be gamed compounds the problem.

There is also one more challenge: everything so far measures a model at a single moment, while production never stops changing.

Offline vs. online evaluation

The most useful way to divide evaluation is not by metric or by method. It is by when the evaluation runs and what data it sees.

Offline evaluation scores a candidate against a fixed dataset before it reaches users. You control the inputs, have reference answers, and can run the same test repeatedly. This is where benchmarks, golden datasets, and regression suites live.

Online evaluation scores live production responses on traffic you have never seen. Rather than comparing them with reference answers or expected outputs, these scores measure whether each response meets the established standards for quality.

Offline evaluation

Online evaluation

What it sees

A curated, fixed dataset

Real production traffic

When it runs

Before exposure

Continuously, after release

Ground truth

Usually available

Rarely available

What it answers

Is this candidate good enough to try?

Is what you are serving right now still good?

What it cannot tell you

How the model behaves on inputs you did not anticipate

Whether a different candidate would have done better

Both types of evaluation are necessary. Offline evaluation is how you avoid shipping something substandard. Online evaluation is how you find out that a prompt scoring well on 500 curated examples degrades on the long tail of real user phrasing, or that a provider updated a model version underneath your configuration.

The two also feed each other. Failures discovered in production become tomorrow's offline test cases, which is how a golden dataset stops being a snapshot of what you imagined and becomes a record of what broke.

Why LLM evaluation matters before and after launch

Air Canada discovered the cost of an unreliable model when its support chatbot misstated its bereavement fare policy and a tribunal held the airline to it. Hallucinations are a common LLM failure mode, and in customer-facing applications, they can carry financial, reputational, and legal consequences. With regulations such as the EU AI Act's transparency provisions and California's SB 53 now in effect, AI accountability is becoming a compliance concern as well as an engineering one.

However, avoiding failures is only one purpose of evaluation. Before deployment, you also need to determine which model is best suited to the task and whether it can meet the application's practical requirements. A large number of models are available for various tasks, and a single model rarely fits every use case, so teams have to compare options rather than default to one. The best model for a given use case also has to satisfy cost, privacy, and latency requirements, not just quality ones.

Evaluation is how you compare candidates fairly against both task performance and operational constraints. It matters again after fine-tuning, because a model adapted to your data has to be rechecked against the downstream tasks you care about.

Model evaluation vs. system evaluation

Model and system evaluation describe what you are testing. Offline and online describe when and where you test it. And the two dimensions can be combined: you can evaluate a model offline against a benchmark, or evaluate a whole system online against live traffic.

Model evaluation

Model evaluation can take several forms.

Intrinsic, or model-level, evaluation examines properties of the model itself using measures such as calibration, or how closely a model's confidence matches its actual accuracy.

Extrinsic, or task-based, evaluation tests how well the model performs on downstream tasks such as translation, summarization, or dialogue.

Teams also conduct behavioral evaluations, which assess how a model behaves in scenarios that require safety, fairness, and robustness. Teams may test whether the model follows instructions consistently, avoids harmful or biased outputs, handles ambiguous or adversarial prompts, protects sensitive information, and remains reliable when inputs are incomplete, misleading, or phrased in unexpected ways. These evaluations can reveal failure patterns that aggregate performance metrics may overlook.

System evaluation

System evaluation determines how well a system performs for a particular use case, such as a RAG-based chatbot. Rather than testing the model in isolation, it evaluates the model within its operating context.

A model that performs well on a benchmark, for instance, can still fail as part of a larger system. For example, the retrieval layer may surface the wrong document, the prompt template may truncate important context, or a tool call may just time out. In any case, when production quality declines, the first useful question is not whether the model changed, but which part of the system failed.

Benchmarks for model evaluation

Because LLM outputs can be evaluated in different ways, model benchmarks are usually best understood by how they score a response: against a known answer, through human preference, or with another LLM acting as the judge.

Automatically scored benchmarks

These benchmarks compare outputs with known answers or verify them through executable tests:

  • MMLU (Massive Multitask Language Understanding): Tests knowledge and reasoning through multiple-choice questions across 57 subjects.
  • GSM8K (Grade School Math 8K): Measures multistep mathematical reasoning using more than 8,500 grade-school-level word problems.
  • ARC (AI2 Reasoning Challenge): Tests reasoning through multiple-choice science questions based on grade-school curricula.
  • TruthfulQA: Measures whether a model provides truthful answers rather than repeating common misconceptions.
  • HumanEval: Evaluates generated Python code by running it against unit tests.
  • SWE-bench: Evaluates whether a model can resolve real-world software issues by applying its proposed patches and running repository tests.

Although people create and validate these datasets, model responses are generally scored automatically rather than reviewed individually by human judges.

Human-preference benchmarks

Chatbot Arena evaluates open-ended responses through crowdsourced, pairwise comparisons. Users submit a prompt, review responses from two anonymous models, and vote for the stronger response.

LLM-judged benchmarks

Other benchmarks use an LLM as the evaluator:

  • AlpacaEval 2.0: Uses an LLM judge to compare model responses with reference responses and reports a length-controlled win rate.
  • MT-Bench: Uses an LLM judge to score responses across multi-turn conversations.

Human- and LLM-judged benchmarks are most useful for open-ended outputs that lack a single correct answer. They can capture qualities such as helpfulness, clarity, and instruction following, but their results may also reflect preferences or biases in the human voters or judge model.

Limitations of model benchmarks

"All models are wrong, but some are useful" applies to benchmarks too. Rely on them, but know what they hide:

Data leakage: A model may have already been trained on the test data, because models ingest enormous quantities of publicly available data that can include some or all of a benchmark.

Overfitting: The pursuit of a benchmark score can produce a model tuned to that benchmark rather than to the underlying capability.

Language and culture bias: Results for English tasks are typically better than those for other languages, such as Arabic or Persian. The heavier weighting of Western sources also introduces cultural bias. If you ask a model how to celebrate an upcoming birthday, the answer will lean toward cakes, balloons, and candles more often than toward any traditions from non-Western cultures.

Benchmark limitations are an active research area. Work on null models has shown that trivial or constant-response models can score surprisingly well on some benchmarks, which says more about the benchmark than the model.

With the exception of Chatbot Arena, which collects live human preferences, the benchmarks above are offline instruments. How well any score carries over depends on how closely your traffic resembles the dataset it was measured on.

LLM evaluation metrics

LLM evaluation rarely comes down to a single score. Teams typically combine several types of metrics because each captures a different aspect of quality: similarity to a reference answer, semantic meaning, factual grounding, safety, or successful task completion.

The right mix depends on your application and on whether the output has a clearly verifiable answer.

Reference-based metrics

Reference-based metrics compare a generated response with one or more expected answers. They work best for tasks such as translation, summarization, and question answering, where representative reference outputs are available.

  • Exact match: Checks whether the generated answer matches the reference exactly. It is useful for questions with a single verifiable answer but too rigid for open-ended language, where several phrasings may be equally correct.
  • METEOR: Compares generated and reference text while accounting for exact word matches, word stems, and synonyms. This makes it more flexible than metrics based entirely on literal word overlap.
  • BERTScore: Uses contextual embeddings to compare individual tokens in the generated and reference texts. It can recognize semantic similarity even when the two responses use different wording.
  • MoverScore: Uses contextualized embeddings to measure how much semantic movement would be required to align the generated text with the reference. Like BERTScore, it focuses more on meaning than exact phrasing.

Reference-based scores are useful, but similarity does not necessarily establish correctness. A response can resemble the reference while introducing a factual error, omitting an important qualification, or failing to satisfy the user's actual request.

Model-level metrics

Some metrics evaluate properties of the language model rather than the quality of a specific application response.

  • Perplexity: Measures how uncertain a model is when predicting a sequence of text. Lower perplexity means the observed text was more predictable to the model. It can help compare language modeling performance, but it does not directly measure factual accuracy, helpfulness, or task completion.

Perplexity is most useful as a model-level diagnostic measurement rather than a complete gauge of response quality.

Grounding and response-quality metrics

For applications such as RAG systems and customer-facing assistants, evaluation must also determine whether a response is supported, relevant, and correct.

  • Faithfulness: Measures whether claims in the response are supported by the information initially supplied to the model, such as retrieved documents or tool results. It helps identify responses that sound plausible but introduce unsupported claims.
  • Factual correctness: Assesses whether the claims in a response are accurate. In this case, correctness may require comparison with external facts rather than only the provided context.
  • Answer relevance: Measures how directly and completely the response addresses the user's question, capturing both whether it leaves out information that would have been useful and whether it pads the answer with information that isn't.
  • Context relevance: Assesses whether the information retrieved or supplied to the model is useful for answering the question. This can help distinguish a generation failure from a retrieval failure.
  • Hallucination detection: Identifies fabricated, unsupported, or contradictory claims. Depending on the application, this may combine checks for faithfulness, factual correctness, citation accuracy, and consistency.

These metrics overlap, but they answer different questions. For instance, a response can be faithful to an inaccurate source, factually correct but irrelevant to the question, or relevant while omitting critical information. Together, they provide a broader measure of how useful the response is to the end user.

Behavioral and safety metrics

Behavioral metrics evaluate whether a model operates within the boundaries expected for its use case.

  • Toxicity: Estimates whether a response contains abusive, threatening, hateful, or otherwise harmful language.
  • Bias and fairness: Examines whether model behavior or outcomes vary unfairly across demographic groups, identities, dialects, or other relevant categories.
  • Robustness: Measures whether the system maintains appropriate performance when prompts are ambiguous, adversarial, misspelled, incomplete, or phrased in unfamiliar ways.
  • Policy compliance: Determines whether the response follows application-specific rules, safety requirements, and restrictions.

These evaluations often require carefully designed test cases because a single aggregate score can obscure severe failures affecting a small group of prompts or users.

Task and system metrics

For applications that take actions rather than only generate text, the most important question may be whether the system completed the task successfully.

  • Task completion: Measures whether the system fulfilled the user's request and reached the intended outcome.
  • Tool selection: Evaluates whether an agent chose the appropriate tool, API, or data source for the task.
  • Tool-call accuracy: Checks whether the system invoked the tool correctly, supplying well-formed arguments that conform to the tool's schema and reflect the user's request.
  • Groundedness of tool results: Evaluates whether the system interpreted the returned information correctly and reflected it faithfully in its final response.
  • Pass@k: Estimates the likelihood that at least one of k generated attempts produces a correct result, such as code that passes the required tests. It's commonly used for code generation and other verifiable tasks.
  • Efficiency metrics: Track operational qualities such as latency, token consumption, cost, and the number of steps or tool calls required to complete a task.

Using metrics offline and online

Many of these metrics can support both offline and online evaluation, but the evidence available changes. Offline evaluation usually applies metrics to a controlled dataset containing reference answers, labels, test cases, or expected behaviors. Online evaluation applies quality criteria to real production interactions, where reference answers are usually unavailable.

As a result, online evaluation often relies on rubrics, classifiers, LLM judges, user feedback, behavioral signals, and task outcomes. Instead of asking only whether a response matched an expected answer, it asks whether the system's behavior met the quality, safety, and performance standards established for the application.

LLM evaluation methods

Three methods cover most practical evaluation work.

LLM-as-a-judge

One of the most common methods is to use an LLM to assess another model's output. You provide the judge with a prompt containing the outputs to compare, and ask it to return a score, a rationale, or both.

For example, a support assistant might answer a customer's question using help docs as a reference. An LLM judge could then assess whether the response is accurate, grounded in the source material, relevant to the question, and clear enough for the customer to act on. Breaking the rubric into distinct criteria produces more useful results than asking whether the response is simply good enough.

LLM judges can evaluate large volumes of responses faster and more consistently than a human doing it manually. You can define scoring guides for qualities such as correctness, relevance, tone, policy compliance, or task completion. This makes them useful for both offline testing and ongoing production monitoring.

Of course, their scores are still model judgments rather than objective truths. Judge models can favor certain response lengths, styles, or phrasings, and their conclusions depend heavily on the scoring guide and prompt. Teams should validate judges against human ratings and use deterministic checks whenever an answer can be verified directly.

Hybrid evaluation

Hybrid evaluation combines LLM judges with human reviewers because automated scoring can miss context, misinterpret unusual cases, or apply the scoring guide inconsistently. Teams will typically use LLM judges to review large volumes of output, then have people check a sample to make sure the results still match the intended quality standards. Human review is especially important for sensitive, subjective, or high-impact decisions.

Red teaming and robustness testing

LLMs are improving quickly, but safety challenges, especially around sensitive data and access permissions, persist, such as:

  • Generating toxic or abusive content
  • Leaking private data if it has been memorized
  • Providing dangerous instructions

Systems that use tools or external data introduce additional risks, including prompt injection, unauthorized actions, insecure tool use, and failures that spread across multiple components.

Safety instructions can be bypassed through jailbreaking, which uses adversarial prompts to circumvent restrictions. These issues are rising at the enterprise level, where AI adoption is quickly outpacing its governance and sensitive data is abundant.

Red teaming attempts to address these issues. Borrowed by cybersecurity from military practice, where a "red team" plays the adversary, red teaming simulates attacks to find where defenses fail. Leading AI labs maintain dedicated red-teaming programs. OpenAI conducts internal, external, and automated red teaming; Anthropic operates a Frontier Red Team; and Google uses an internal AI Red Team to test AI systems for security, privacy, and abuse risks.

Why LLM judges need to be evaluated too

If evaluation is going to drive decisions in production, the LLM judge must be tested for strengths and weaknesses just like any other system component. Its reliability should be treated as a quality that needs to be continuously measured and managed, not just assumed.

Research from RAND published in March 2026 tested four frontier judges across four benchmarks and found none were uniformly reliable. On the paraphrase test, where responses were reworded but their meaning preserved, the weakest judge matched the expected score only 40% of the time. Later work from July 2026 went further, finding that a strong judge flipped roughly 15% of its verdicts when the two answers were shown in reverse order. Stacking similar models into a jury helped less than expected because their errors tend to correlate.

The conclusion is not that LLM judges are useless. It means their scores contain some uncertainty. Teams should verify whether judgments remain consistent, compare them with human reviews, and avoid basing important decisions on a single score or evaluation run.

Building reliable LLM judges

How you design your evaluation process is just as important as the model you choose. A few practices help make judge scores more stable and useful in production.

Aggregate before you act. A judge may be inconsistent on a single response, but still reveal reliable patterns across hundreds of them. Before an alert or automated action can trigger, define how many responses must be scored and over what period for the result to be statistically significant. Otherwise, one unusual result can cause the system to react to noise instead of a real quality problem.

Use more than one signal. A single composite score hides failure modes. A response can score in the acceptable range while containing fabricated content, because the fabrication seemed knowledgeable and on topic. Pair judge scores with deterministic checks, such as schema validity, citation presence, and retrieval hit rate, so different failure types have different detectors.

Periodically validate the judge against humans. Agreement between judge verdicts and human spot-checks can drift, particularly in specialized domains. Don't carry a single "judge vs. human" disagreement percentage from one evaluation task to another and assume it means the same thing. Different tasks have different levels of subjectivity, and even human reviewers often disagree on subjective quality, so the baseline level of agreement changes from task to task. Establish a baseline through an initial calibration run, then recalibrate when agreement consistently falls outside that range rather than relying on a single threshold.

Version the judge like any other dependency. A judge model that silently updates will change your measurement without changing your system. When your scores move, the first question is whether the thing measuring them moved.

A practical LLM evaluation example

The following example uses a text summarization task on the CNN/DailyMail dataset, with LLM-as-a-judge to determine which of two models performs better. Two models generate summaries, and a third model judges them.

Define the judge function

Define a comparison function that takes two summaries and returns a verdict. State the criteria explicitly rather than asking for a general preference. "Clearer, more faithful to the source, and free of information not present in the article" produces more consistent verdicts than "better."

Run the model comparison

Make both models write summaries under the same limits (for example, the same length) so the comparison is fair. Keep a running win/loss score for each model, and store both summaries next to the approved human summary you're using as the reference. After each batch, add up the scores and keep tracking results over time.

Check the judge's decision

Run the comparison function on each pair of model outputs and record the judge's verdict.

Two safeguards will make the results more reliable. First, present the responses in both orders to check for position bias, which is the tendency of an LLM judge to favor an answer for its placement (such as being first) rather than its quality. Count the result only when the judge selects the same response both times and flag any inconsistent decisions for a separate review.

Second, make sure to use a judge from a different model family than any of the candidates when possible, which will reduce the chance that it will favor outputs resembling its own style or behavior.

You'll then have completed the offline comparison for the selected dataset. And while it will show which candidate performed better under these given test conditions, remember it will not establish how either model will perform on new requests or live production traffic.

What to do when an evaluation fails in production

Most evaluation guidance ends at "monitor performance in production, and alert when quality drops." But alerting is where the problem starts, not where it ends.

The harder part is identifying what changed, limiting the impact, correcting the problem, and confirming that the fix worked.

A reliable production evaluation loop typically includes five stages:

Stage

What happens

What it requires

Detect

A judge score, error rate, or latency metric crosses a threshold on live traffic

Online evaluation running continuously at a known sampling rate

Attribute

The regression is tied to a specific change: this prompt version, this model, this configuration

Every response carries the identity of the configuration that produced it

Bound

Exposure is limited to the cohort already receiving the change

The change was released progressively, not to everyone at once

Act

Traffic returns to the last known good configuration, or escalates to a more capable one

The configuration can change without a redeploy

Verify

Scores return to baseline, and the failing case enters the offline dataset

The loop feeds back into preproduction testing

In production, a response is shaped by more than just the model. The prompt version, model version, retrieval settings, tools, temperature, and other configuration choices can all affect the result. Attribution means recording which combination produced each response.

Attribution fails when responses are not tagged with the configuration that generated them. If you cannot tell which prompt version produced a bad answer, a quality drop becomes a research project. If every response carries its configuration identity, it is a lookup.

Action fails when the only way to change model behavior is to deploy. If reverting a prompt requires a pull request, a build, and a release, your recovery time is your deployment pipeline's cycle time, no matter how fast your evaluation detected the problem. Detection speed only matters if the response can keep pace with it.

Reverting a configuration only helps when the configuration caused the regression. If quality dropped because a provider deprecated a model version, or because your retrieval corpus drifted, rolling back to the last known good variation restores a setting that is no longer good. Automated action shortens recovery for the failures you caused, but it cannot diagnose the ones you inherited.

This is the distinction between observability and control: LLM observability tells you what is happening and why, which opens the loop, but something still has to close it.

Evaluation as a runtime control signal with LaunchDarkly

Evaluation is most useful when teams can act on the results without rewriting or redeploying the application. LaunchDarkly AgentControl treats prompts, models, and parameters as configuration rather than code, which is what makes an evaluation score capable of changing production behavior.

AgentControl AI Configs store model settings, prompts, tool definitions, and parameters as versioned variations, with version control and audit trails on every change.

A variation bundles the model, prompt, and parameters into a versioned configuration, so switching variations changes the complete configuration atomically rather than mixing independently managed settings. This is the same mechanism as a feature flag, the unit of runtime control that lets teams change behavior in production without changing code, now applied to prompts, models, and parameters. Configuration changes propagate globally in less than 200 ms, without requiring a redeploy.

Offline Evaluations let teams compare AgentControl variations before exposing them to production traffic. Using the LaunchDarkly playground and evaluation workflow, teams can run prompts from a reusable dataset through different combinations of models, instructions, and parameters, and then attach criteria or LLM judges to score the outputs. Because datasets, criteria, and custom judges can be reused, teams can apply the same quality standards across repeated tests and projects. The results help determine which variation should advance toward release, while still reflecting performance on controlled test cases rather than real user behavior.

Online Evaluations extend the process into production by attaching judges directly to AgentControl AI Configs variations. The LaunchDarkly built-in judges score responses for accuracy, relevance, and toxicity. Custom judges let teams define their own criteria for domain-specific requirements. Those scores are recorded against the variation that produced each response, so teams can compare quality alongside the AI metrics LaunchDarkly already tracks, including latency, cost, and user satisfaction. Teams can also set the percentage of production responses evaluated, balancing coverage against the additional cost and latency of judge calls.

That sampling rate should be paired with a clearly defined measurement window. A quality threshold based on only a few judged responses may reflect normal judge inconsistency rather than a meaningful regression. Teams should decide how many evaluations must accumulate, how long the window should remain open, and whether a decline must persist before LaunchDarkly is allowed to pause, roll back, or switch a configuration.

With that in place, LaunchDarkly can connect evaluation and operational metrics to several forms of runtime control:

  • Guarded Rollouts monitor a release against your chosen metrics while exposure ramps. Rather than firing on a single threshold crossing, LaunchDarkly applies sequential testing to detect statistically significant regressions against the tolerance you define. When one is detected, the rollout pauses, and where automatic rollback has been enabled, the release reverts without waiting for a person.
  • Variation switching returns traffic to the last known good configuration in one change, propagating in less than 200 ms, not a deploy cycle.
  • Targeting rules bound exposure before anything goes wrong. Internal users come first, then a percentage of traffic, then a segment, so a bad variation reaches a cohort rather than a customer base.
  • Adaptive Triggers automatically switch to a fallback configuration when a monitored metric breaches a threshold.

This runs on the LaunchDarkly controlled release infrastructure, extended into a control surface built for runtime AI configuration. A single AgentControl config variation bundles the model, prompt, parameters, and tools, so those elements can be targeted by segment, measured per variation, released to a percentage of traffic, and changed at runtime without a redeploy. Evaluation scores feed the same gating layer: Online Evaluations attach accuracy, relevance, and toxicity judges to a variation, and Guarded Rollouts use those scores to pause exposure or roll back automatically when a threshold is violated.

Automated action is only as good as the statistic behind it. A threshold set on too few samples will fluctuate too much, and a rollback that fires on an unreliable judge may be unnecessarily costly. Runtime control shortens the distance between a score and a response, but it cannot decide what the score should have been.

7 LLM evaluation best practices

1. Align evaluation to real-world use

Benchmarks like MMLU are useful but may not reflect your specific task. A model that excels at trivia can still fail at customer support. Test on data close to your actual use case.

2. Give metrics the right context

Combine surface-level metrics with semantic checks and human review, because a chatbot might get the wording right and still misread the user's intent.

3. Re-evaluate as you go

Model updates shift performance. Test after fine-tuning, after retraining, and after a provider version change. One good evaluation run is a snapshot, but consistent testing is still necessary to maintain optimum performance.

4. Use both general and custom tests

Benchmarks such as HELM or TruthfulQA demonstrate general capability, but domain-specific sets built from your own data, whether those are legal clauses, customer emails, or support transcripts, will be the true evaluation of what you ship.

5. Balance scale and judgment

Automated scoring, including LLM judges, lets you evaluate a lot of outputs quickly. Human review is how you check that those automated scores still match what you consider "good." Regularly have people review a sample of the outputs the judge scored. If the judge's scores and human ratings start to disagree more often over time, treat that as a sign the judge needs to be re-validated or updated.

6. Track quality and efficiency together

High accuracy is not enough if latency triples or costs spike. Log token counts, p95 latency, and cost per successful answer rather than cost per call.

7. Close the loop in production

Deploy continuous checks on live traffic, and define what happens when they fail before they fail. Decide the threshold, the sample window, the action, and the fallback configuration in advance, so the response to a quality drop is a configuration change rather than an investigation.

From scoring models to acting on scores

Evaluation is how you find out whether a model is good enough. That has always been true, and the benchmarks, metrics, and methods above are how you do it well.

What has changed is where the work ends. When behavior is probabilistic and the model underneath you can shift without notice, a score that cleared a threshold before release is evidence about the past. The teams that stay reliable keep scoring after release, and wire those scores to something that can be acted upon.

Hireology built exactly that loop on LaunchDarkly. The hiring platform runs generative AI across industries as varied as healthcare, retail, and automotive, where the same prompt can return 20 different answers and several hallucinations. So instead of trusting one offline check, its system scores models and configurations against its own quality criteria as part of the release process. The team can test three verticals with 10 tests each in less than 13 seconds. Instead of grading models once, Hireology can rescore on demand and see immediately where quality moves, while the team still decides what ships.

Bring your evaluation scores into production decisions with LaunchDarkly AgentControl.

FAQs

What is LLM evaluation?

LLM evaluation measures how well a large language model performs across tasks, contexts, and safety dimensions. You can evaluate a model on its own, or evaluate a whole system such as a RAG chatbot in its operating context.

What is the difference between offline and online LLM evaluation?

Offline evaluation scores a candidate against a fixed dataset before release. Online evaluation scores live production responses continuously, on traffic your test set never contained. Offline keeps you from shipping something bad; online tells you whether what you are serving right now is still good.

What are the main LLM evaluation metrics?

No single score is enough, so teams combine types: reference-based metrics such as exact match and BERTScore, model-level metrics such as perplexity, grounding metrics such as faithfulness and answer relevance, safety metrics such as toxicity and bias, and task metrics such as task completion, pass@k, latency, and cost.

What is LLM-as-a-judge?

LLM-as-a-judge uses a capable LLM to score another model's output against a rubric. It handles far more responses than manual review and works both offline and in production.

How do you build a reliable LLM judge?

Four practices can help:

  • Aggregate scores across many responses before acting, so you react to regressions and not noise.
  • Pair judge scores with deterministic checks like schema validity and citation presence.
  • Spot-check the judge against human reviewers and recalibrate when agreement drifts.
  • Version the judge like any other dependency, so you know whether your scores moved or your ruler did.

What are the most common LLM benchmarks?

MMLU, GSM8K, ARC, and TruthfulQA score answers automatically, while HumanEval and SWE-bench run generated code against tests. Chatbot Arena uses human preference voting, and AlpacaEval 2.0 and MT-Bench use an LLM judge.

What are the best practices for evaluating an LLM?

Test on data close to your real task, combine metrics with human review, re-evaluate after every fine-tune or provider change, use both public benchmarks and your own domain sets, track cost and latency alongside quality, and decide in advance what happens when production checks fail.

How does LaunchDarkly turn evaluation scores into runtime action?

LaunchDarkly AgentControl stores prompts, models, and parameters as versioned configuration, so a score can change production behavior without a redeploy. Offline Evaluations compare variations before release, Online Evaluations attach judges to live traffic, and Guarded Rollouts and Adaptive Triggers (closed beta as of August 2026) use those scores to pause, roll back, or switch configurations automatically.

Like what you read?
Flaming pointer icon
See what's right for you.

Get started in minutes and scale seamlessly as your project grows.

Letter in envelope icon
Sign up for our newsletter

Get all the content, tips, and news you can use.

By supplying my contact information, I authorize LaunchDarkly to contact me with personalized marketing communications about our products and services. See our Privacy Policy for more details, or Opt-Out at any time.