Case Study · 7 min read
A Router Trained on 63 Product Listings Matched Claude Sonnet 5 at 4% of the Cost
Four models localized 63 listings into seven languages. A router trained on that scoreboard then scored 0.909 on 28 unseen listings, Sonnet 0.911, at $0.00033 versus $0.00885 per listing.
| Option | Panel score | Flawless | $ / listing | Speed |
|---|---|---|---|---|
Claude Sonnet 5, alwaysclaude-sonnet-5 ★ Highest score by 0.002. 12 of 28 listings flawless, the most of the three. 95% interval 0.88 to 0.94. | 0.911 | 12 / 28 | $0.00885 | 6.8s |
Trained routerrouter Sent 14 listings to gpt-5.6-luna and 14 to DeepSeek V4 Flash. It had never seen these 28 listings. 95% interval 0.89 to 0.93. | 0.909 | 7 / 28 | $0.00033 | 18.4s |
gpt-5.6-luna, alwaysgpt-5.6-luna Fastest of the three. 95% interval 0.87 to 0.92. | 0.901 | 7 / 28 | $0.00046 | 4.8s |
How it works
The router picks a model per listing from the 63 scored examples; the held-out 28 were never in its training set.
What we actually did
We took the existing Listing Localization flow: an English product listing and a target language go in, a localized title, description and bullets come out.
We generated 91 English listings across 13 product categories and assigned each a target language, 13 per language: Spanish, German, French, Portuguese, Hebrew, Japanese and Arabic.
63 listings were the training run and 28 were held out. The held-out listings were never shown to the router before the final run.
Four models localized the 63 training listings: Claude Sonnet 5, gpt-5.6-luna, Gemini 3.7 Flash and DeepSeek V4 Flash. Three judges from three vendors scored every output blind on four criteria, a quarter point each: meaning fidelity, native fluency, keyword adaptation and structure.
A router was then trained on that scoreboard and run on the 28 held-out listings next to Sonnet and luna.
The training run
| Model | Panel score | 95% interval | Flawless | $ / listing | Speed |
|---|---|---|---|---|---|
| Claude Sonnet 5 | 0.914 | 0.90 to 0.93 | 27 / 63 | $0.00906 | 7.3s |
| gpt-5.6-luna | 0.908 | 0.89 to 0.93 | 24 / 63 | $0.00045 | 4.9s |
| Gemini 3.7 Flash | 0.897 | 0.88 to 0.92 | 16 / 63 | $0.00328 | 4.1s |
| DeepSeek V4 Flash | 0.893 | 0.86 to 0.92 | 16 / 63 | $0.00021 | 13.7s |
63 listings per model, 252 localizations, 756 panel judgments. Flawless = panel score of 0.95 or higher.
Sonnet cost 20 times luna per listing and scored 0.006 higher. The intervals of all four models overlap.
Per language. The table below is the mean panel score per model per target language, nine listings each.
| Model | Spanish | German | French | Portuguese | Hebrew | Japanese | Arabic |
|---|---|---|---|---|---|---|---|
| Claude Sonnet 5 | 0.96 | 0.96 | 0.94 | 0.94 | 0.82 | 0.89 | 0.90 |
| gpt-5.6-luna | 0.96 | 0.94 | 0.94 | 0.91 | 0.86 | 0.85 | 0.89 |
| Gemini 3.7 Flash | 0.95 | 0.89 | 0.95 | 0.92 | 0.85 | 0.88 | 0.84 |
| DeepSeek V4 Flash | 0.92 | 0.94 | 0.93 | 0.83 | 0.82 | 0.93 | 0.89 |
Every model scored lowest on Hebrew. On Hebrew, luna scored 0.86 and Sonnet 0.82; on Japanese, Sonnet 0.89 and luna 0.85; on Arabic, Sonnet 0.90 and luna 0.89.
DeepSeek V4 Flash scored 0.94 on German and 0.93 on Japanese at $0.0002 per listing, and 0.83 on Portuguese.
What the router learned
The router is the KNN router from UIUC’s open-source LLMRouter library. Each listing is embedded; its nearest training listings vote for the model with the best score minus cost weight times cost.
Evaligo evaluates the router leave-one-out: each of the 63 listings is routed using the other 62, and the measured score and cost of the chosen model are recorded. Training and evaluation took 19 seconds.
| Working point | Score | $ / listing | Where listings went |
|---|---|---|---|
| Quality only (cost weight 0, K = 7) | 0.914 | $0.00906 | Sonnet for nearly all |
| Cost weight 0.02 (K = 7) | 0.910 | $0.00849 | mostly Sonnet |
| Cost weight 0.05 to 1.0 (K = 7) | 0.908 | $0.00045 | luna for all 63 |
| Cost weight 1.5 and above (K = 7) | 0.893 | $0.00021 | DeepSeek V4 Flash for all 63 |
Leave-one-out on the 63 training listings. The error bar at every point is about ±1.8 points; the router’s interval overlaps Sonnet’s at every working point.
There was no mixed working point that beat both Sonnet and luna. Once cost carried any weight, the router sent every training listing to luna.
We saved the working point at cost weight 0.5 with K = 7 and ran the router on the 28 listings it had not seen.
The held-out run
The scoreboard at the top of this post is that run: Sonnet always, luna always, and the router.
On unseen listings the router did not send everything to luna. It sent 14 to luna and 14 to DeepSeek V4 Flash. Luna and DeepSeek were within 0.015 of each other in the router’s utility, so the neighbour votes split on new inputs.
| Held-out, 4 listings each | Spanish | German | French | Portuguese | Hebrew | Japanese | Arabic |
|---|---|---|---|---|---|---|---|
| Claude Sonnet 5 | 0.96 | 0.94 | 0.88 | 0.97 | 0.90 | 0.92 | 0.80 |
| Router | 0.94 | 0.93 | 0.94 | 0.93 | 0.85 | 0.92 | 0.84 |
| gpt-5.6-luna | 0.93 | 0.92 | 0.94 | 0.90 | 0.83 | 0.93 | 0.85 |
| Router sent to | 3 DeepSeek, 1 luna | 1 DeepSeek, 3 luna | 2 and 2 | 2 and 2 | 2 and 2 | 1 DeepSeek, 3 luna | 3 DeepSeek, 1 luna |
Four listings per language per option; individual cells carry wide error bars.
Router: 0.909 at $0.00033 per listing. Sonnet: 0.911 at $0.00885. Luna: 0.901 at $0.00046. Per judge, the router scored 0.897 (gpt-5.4-mini), 0.831 (Kimi K2) and 0.980 (GLM-5); Sonnet 0.868, 0.850 and 0.976.
The router’s 18.4 seconds per listing is the average of luna’s 4.8 seconds and DeepSeek’s 13.7 seconds plus the routing decision, which takes about 0.2 seconds once the encoder is loaded.
Where the judges took points
Across the 252 panel judgments of the held-out run, keyword adaptation was named in 182: titles and bullets translated word for word instead of using the terms buyers search for in that market.
Awkward or unnatural phrasing was named 62 times, wrong or mixed script 45 times, literal translation 24 times, added detail 14 times.
“Sérum Visage Vitamine C is slightly awkward compared to the more natural French e-commerce term.”
Kimi K2 on a Sonnet output. Kimi K2 gave the lowest scores of the three judges to every model; GLM-5 gave the highest.
The fine print
Data: 91 English listings generated by gpt-5.6-luna across 13 product categories, each with a title, a 60 to 90 word description and five bullets. Languages assigned round-robin, 13 each; 9 per language for training and 4 for the held-out run.
Models: Claude Sonnet 5 on Anthropic, gpt-5.6-luna on OpenAI, Gemini 3.7 Flash on Google, DeepSeek V4 Flash via OpenRouter; the same prompt, temperature 0.7, 6,000 max tokens, strict JSON schema.
Judges: gpt-5.4-mini, Kimi K2 and GLM-5, each scoring 0 to 1 blind against the four-criterion rubric with a 0.5 deduction for the wrong language; scores averaged. None of the judges is a contestant.
Router: KNN, K = 7, cost weight 0.5, sentence embeddings computed on the Evaligo server; candidates were the four training-run models with their measured scores and costs.
Cost: training run $0.82 in model calls and $4.44 in judging; held-out run $0.27 and $1.89. The whole study, judging included, was $7.42.
Run this on your own use case
Any finished Evaligo experiment with two or more models and judge scores can train a router with one click. The router then appears in the model list and runs in the next experiment like any other model, with error bars and a note on how many more samples would tighten them.
The flow, the dataset, the rubric and both runs above are reproducible in Evaligo.
Ready to Build This?
Start building AI workflows with Evaligo's visual builder. No coding required.
Need Help With Your Use Case?
Every business is different. Tell us about your specific requirements and we'll help you build the perfect workflow.
Get Help Setting This UpFree consultation • We'll review your use case • Personalized recommendations
