Visual AI workflow builder with A/B testing for models and prompts

Build a flow of prompts, agents and tools. Run it on several AI models, score every output with blind AI judges, and compare quality, cost and speed. Then deploy any flow as an API.

Drag-and-drop flow editorAny model, incl. open sourceBlind AI judge panelsQuality, cost and speed per outputModel routing per inputDeploy any flow as an API
Evaligo A/B test scoreboard: six AI models compared on the same 30 inputs, scored by three blind judges, with cost and speed per output

A real Evaligo experiment: six models, the same 30 inputs, three blind judges. Read the full write-up.

What testing found on our own workflows

Three experiments run on Evaligo. Each ran every model on the same inputs and recorded the score and the cost per output.

Language detection · 129 texts

Six models, the same score, 32× the price

Every model scored 0.99 or better on the same 129 texts. The most expensive one cost 32 times the cheapest and did not score higher.

ModelScoreCost
gpt-5.4-mini1.000$0.00015
claude-sonnet-4-61.000$0.00063
claude-haiku-4-50.992$0.00023
gemini-3.5-flash0.992$0.00485

Scored against a gold answer set. Cost is per text.

Listing localization · 63 listings, 7 languages

0.006 apart, 20× apart in cost

Claude Sonnet 5 scored highest, but gpt-5.6-luna came within 0.006 of it for a twentieth of the price per listing.

ModelScoreCost
claude-sonnet-50.914$0.00906
gpt-5.6-luna0.908$0.00045
gemini-3.7-flash0.897$0.00328
deepseek-v4-flash0.893$0.00021

Blind three-judge panel. Cost is per listing.

Read the case study

Writing level · 87 texts

The cheaper model graded better

Flash scored higher than the Pro model of the same family, and cost less per text. Price did not predict quality.

ModelScoreCost
gemini-3.5-flash0.977$0.00648
gemini-3.1-pro0.948$0.00774
gemini-3-flash0.937$0.00175
gemini-2.5-pro0.845$0.01151

Graded against a reference answer set. Cost is per text.

01

Compare models, prompts and flows on the same inputs

Drag-and-drop flow builder

Describe the flow in one sentence and Evaligo builds it: prompts, web scraping, data extraction, tools and agents as connected nodes. Drag nodes on the canvas to edit it, or change it by instruction.

View DocsBuild a flow free
Drag-and-drop flow builder - Evaligo A/B testing
Any model as a variant, including open source - Evaligo A/B testing

Any model as a variant, including open source

Pick the models you want to compare: OpenAI, Google, Anthropic, and open-source models such as DeepSeek, MiniMax, Kimi and Qwen through one OpenRouter key. Each variant runs the same flow on the same inputs; only the model, prompt or reasoning setting changes.

View DocsCompare models free

Test on your own data, or generate it

Run every variant over a dataset of real inputs. If you have none yet, Evaligo generates a realistic test set from a one-line description, so a 30-sample comparison does not wait for you to collect data.

View DocsCreate a test set
Test on your own data, or generate it - Evaligo A/B testing
02

Judge every output blind

Scoring rubrics, written once

Write the criteria once: what a good output must do, and what costs points. AI judges score every output against the rubric without knowing which variant produced it, and record a one-line reason per score.

View DocsWrite your first rubric
Scoring rubrics, written once - Evaligo A/B testing
Judge panels from several vendors - Evaligo A/B testing

Judge panels from several vendors

Use one judge model or a panel of three from different vendors, averaged. Per-judge scores stay visible, so you can see where the judges agree and where they do not.

View DocsSet up a judge panel

Repeat runs

Model output and judging both vary between runs. Re-run an experiment with one click and compare runs side by side before you decide.

View DocsRun an experiment
03

Read the scoreboard and deploy

Quality, cost and speed per output

Every experiment ends in one table: score, flawless outputs, cost per output and latency for each variant, with a score distribution per sample. Export it as a PDF report.

View Docs

Promote the winning variant

Make the winning model, prompt or flow the live version with one click. The flow keeps its inputs, outputs and API endpoint.

View Docs

Train a router on the results

When an experiment finishes, one click trains a router on its scoreboard. It learns which model handles which kind of input, shows quality against cost with 95% error bars so you can choose a working point, and then runs as a model in your next experiment or in a deployed flow.

Read the case study

Deploy as a production API

Any flow deploys as a REST API with authentication, documentation and code examples. Batch processing and monitoring are built in.

View Docs

Benchmarks we published, run on Evaligo

Every number below comes from an experiment you can reproduce: the same flow, the same inputs, blind judges, cost and speed per output.

Model routing

Router: 0.909 vs Sonnet’s 0.911, at 4% of the cost

In the listing localization test, a router trained on the four-model scoreboard picked a model for each of 28 listings it had not seen. It scored 0.909, Claude Sonnet 5 scored 0.911, at $0.00033 versus $0.00885 per listing.

Any finished experiment can train a router. You pick the quality and cost trade-off with 95% error bars, and the router runs in your next experiment or flow like any other model.

Listing localization, 28 held-out listings, judge score and cost per listing
Claude Sonnet 5 (every input)0.911
$0.00885 per listing
Trained router (picks a model per input)0.909
$0.00033 per listing

Bar length is cost. Scores from a blind three-judge panel. Method: KNN router from UIUC’s open-source LLMRouter library.

Pricing: free plan, Pro at $7 a month

Both plans include the flow builder, experiments and API deployment.

Start Here

Free

$0

  • 500 credits/month
  • Visual workflow builder
  • Web scraping & browser automation
  • Models from OpenAI, Anthropic, Google and OpenRouter
  • All integrations (GitHub, Bitbucket, etc.)
  • Deploy as API endpoints
  • One-time scheduled runs
  • Bring your own API keys (unlimited)
Get Started Free

Pro

Cancel Anytime
$7/month

  • 10,000 credits/month
  • Buy more credits when you need them
  • Recurring schedules (hourly, daily, weekly)
  • 10 concurrent API executions
  • Priority email support
Upgrade to Pro

Credits are used for AI model calls with platform API keys. Use your own API keys for unlimited usage on any plan.

Compare models on your own data

Run your prompt or flow on several models, let blind judges score the outputs, and compare them on quality, cost and speed.

6 models
compared in one experiment, open source and closed
3 judges
from three vendors scoring every output blind
30 inputs
the same 30 bios run through every model
Free plan, no credit card.

Free plan • 500 credits a month • No credit card • Cancel anytime