An evaluator is a model that scores one output against one goal and returns a number between 0 and 1 with a one-line reason. A prompt or flow can have several evaluators, one per criterion, and the same evaluator can be saved to your library and reused across prompts, flows and experiments.

What an evaluator consists of

  • Name and an optional description, so it can be found in the library.
  • Goal: the one criterion this evaluator judges, in plain language. For example, “The reply answers the customer’s question without inventing policy details.”
  • Template: the prompt the evaluator model receives. It has three placeholders, filled in at run time.
  • Model and settings: provider, model, temperature, max tokens, top-p, frequency and presence penalties. Each evaluator has its own, so a strict judge and a lenient one can run side by side.

The default template

New evaluators start from this template. {sample_variables} receives the input row,{output} the text being judged, and {goal} the evaluator’s goal. You can edit the wording; keep the three placeholders and the JSON instruction, because Evaligo reads score and reason from the reply.

Default evaluator template
You are evaluating an LLM's response.

Original Sample data:
{sample_variables}

We run a prompt on this data.

LLM's Response:
{output}

Evaluation Goal:
{goal}

Rate the response on a scale from 0.0 to 1.0 based on how well it meets the goal, and provide a brief reason.
Return your output as JSON with 'score' (float 0.0-1.0) and 'reason' (string).
Example: {"score": 0.85, "reason": "The response is accurate and well-structured."}

Save to the library and reuse

  1. 1

    Save In the Playground’s Evaluators tab, each evaluator has a Save to Library action. It stores the name, goal, template, model and settings.

  2. 2

    Load Load from library lists your saved evaluators, newest first, and adds the one you pick to the current prompt as a copy. Editing the copy does not change the library entry.

  3. 3

    Manage The Evaluators page in the app lists the library with how many times each evaluator has been used, and lets you edit or delete entries.

Where evaluator scores are used

  • Playground runs show each evaluator’s score and reason per sample, and the average per evaluator.
  • Evaluation nodes in a flow run evaluators on the output of the node before them, so a flow can score itself while it runs.
  • Experiments can score variants with a judge, with a reference answer, or with the flow’s own evaluation nodes: in that mode the scoreboard shows the average of the in-flow criteria per variant.
Judge models
Judges in experiments, and using a panel of them.
Custom evaluations
Writing goals for your own domain.
Run an experiment
Scoring variants with evaluators, judges or references.