The Cheapest AI Won: We Made 3 Models Rewrite 30 Cringe Dating Bios
We built a Dating Profile Doctor flow from one sentence, fed it 30 deliberately bad bios ("gym, tacos, dogs"; "partner in crime wanted"), and had a blind AI judge score three budget models on authenticity and cringe-avoidance. The cheapest model won on every axis — here is the scoreboard.
Everyone assumes writing quality is what you pay for. So we built a Dating Profile Doctor — a flow that rewrites a dating-app bio in three tones (witty, sincere, direct) plus one honest tip — fed it 30 deliberately bad bios, and had a blind AI judge score three budget models on authenticity, cringe-avoidance, tone-distinctness, and tip quality. The cheapest model won on every axis:
| Model | Authenticity | Perfect 1.0s | $ / profile | Speed |
|---|---|---|---|---|
| gpt-5.6-luna ★ wins on every axis |
0.86 |
17/30 | $0.0006 | 4.6s |
| gemini-3.5-flash 2nd on quality — 27× the price |
0.79 |
7/30 | $0.0163 | 8.4s |
| claude-haiku-4-5 cheap, but flat and generic |
0.67 |
4/30 | $0.0024 | 5.5s |
Authenticity is the average of a 4-criterion blind rubric (authenticity, cringe-avoidance, tone-distinctness, tip quality; −0.25 per failed criterion, −0.5 for invented facts) over the same 30 bios. “Perfect 1.0s” = profiles where nothing failed. $ / profile is flow cost only — judging cost is identical across models and excluded.
Bottom line: gpt-5.6-luna wrote the most human bios — flawless on 57% of profiles — at $0.0006 each, 27× cheaper than gemini-3.5-flash and 4× cheaper than claude-haiku-4-5. Paying more bought nothing here.
Before and after: what a 1.0 looks like
Before (real test bio)
“Partner in crime wanted. I love to laugh, live for spontaneous adventures, and am fluent in sarcasm. Swipe right if you can keep up!”
After — luna's sincere version
“I’m happiest at a live show, by the beach, or laughing through a trivia night. I love craft cocktails, reality TV, and making overly ambitious dinner plans — and I’m looking for someone kind, funny, and emotionally available to share them with.”
And the honest tip
“Consider removing ‘partner in crime’ and ‘swipe right if you can keep up’ — they’re common phrases that can sound generic or slightly competitive; your specific interests do more to show your personality.”
4 things we learned
Price doesn't predict writing quality.
The cheapest model per profile scored highest on human-ness. The most expensive one (flash) was second — and most of its cost was invisible thinking tokens, not better words.
“Cringe” is measurable.
A rubric that names the failure modes (clichés, forced jokes, invented facts, blurred tones) turns a vibe into a number a blind judge scores consistently — and models differ on it far more than on grammar.
Small models fail on the hardest inputs.
Haiku's three 0.25 scores all came from cliché-stuffed bios — it echoed the clichés back instead of replacing them. Luna's worst case was tones blurring together; it never invented facts.
30 samples changed the ranking.
At 6 samples, flash and haiku tied. At 30, flash pulled clearly ahead of haiku — and luna's lead widened. If you're picking a model on 5 test runs, you're guessing.
How it works: rewrite → judge → scoreboard
The fine print
- Inputs: 30 synthetic bios generated to be deliberately mixed — near-empty (“gym, tacos, dogs”), cliché-stuffed, oversharing, bitter, one all-emoji — across ages 19–62 and orientations.
- Judge: gpt-5.4-mini (not a contestant), scoring each output 0–1 against the rubric without knowing which model wrote it.
- Sample: 30 profiles × 3 models = 90 runs, 0 failures. Enough to rank clearly; treat the second decimal as noise.
- Total experiment cost: about $0.90 including all judging.
Benchmark your own use case
Describe a flow in one sentence, run every candidate model against your real inputs, and let a blind judge pick the winner before you commit — that is what Evaligo experiments are for. The same setup works for support replies, product copy, summaries, or anything else where “sounds human” matters more than a leaderboard rank.
Ready to Build This?
Start building AI workflows with Evaligo's visual builder. No coding required.
Need Help With Your Use Case?
Every business is different. Tell us about your specific requirements and we'll help you build the perfect workflow.
Get Help Setting This UpFree consultation • We'll review your use case • Personalized recommendations