Case Study5 min read

The Expensive AI Apologized Its Way Out of a Raise — Then Its Newest Version Fixed It

We had three budget AI models negotiate 30 job offers (nurse to staff engineer, hourly to $200k+), scored blind by an AI judge. gemini-3.5-flash finished last for apologizing too much — then 3.7-flash fixed the tone and jumped to second. The cheapest model still won every axis.

By Danny Lev, Founder & CEO

Can a cheap AI coach a real salary negotiation? Thirty job offers, three models, one blind judge — run on the newest version of each model. The cheapest one won:

Model Negotiation score Plans 0.95+ $ / plan
gpt-5.6-luna
winner on every axis, again
0.82
14 / 30 $0.0012
gemini-3.7-flash
big upgrade over 3.5 — and cheaper
0.75
8 / 30 $0.0095
claude-haiku-4-5
grounded, but the asks land weak
0.68
3 / 30 $0.0055

30 offers per model, judged blind by a separate AI. Model versions as of Aug 24, 2026. “0.95+” = nothing meaningful to deduct.

We first ran this with gemini-3.5-flash. It finished last — its scripts kept apologizing. The newest 3.7-flash fixed the tone, climbed to second, and costs 3.5× less. Versions matter as much as vendors.

What we did

We built a Salary Negotiation Coach: give it the offer, your experience, your leverage, and your biggest worry. It writes the negotiation email, a 60-second phone script, replies to the three most likely objections, and a realistic ask.

The 30 test scenarios ranged wide: a nurse, a teacher, a retail shift manager, a staff engineer. Hourly wages and $200k+ packages. Half had no competing offer. Worries included “I hate conflict”, “I’m on a visa”, and “I was told the band is fixed”.

A separate model judged each plan blind on four things: grounded (no invented market data — automatic −0.5), assertive-but-not-adversarial, specific to this person’s leverage, and a defensible ask with a fallback.

A real output

The scenario

Senior backend engineer, 7 years, offered $128k at a fintech startup. One competing offer at $135k. Worry: “I don’t want to seem greedy and lose the offer — I actually like this team.”

Luna’s one tip

“Frame the conversation as wanting to join while asking the company to close the gap with your concrete $135,000 alternative offer — not as a demand or a threat.”

And its reply to “the offer already beats your current salary”

“I appreciate that, and I’m grateful for the offer. My request is based not only on the increase from my current salary, but also on the scope of my experience...”

What we learned

1. Versions matter as much as vendors.
Swapping gemini-3.5-flash for 3.7-flash — nothing else changed — moved its score from 0.68 to 0.75, cured the apologetic tone, and cut its price 3.5×. If you benchmarked a vendor six months ago, that result is stale.

2. Tone failures are directional.
Old flash failed soft — polite surrender. Haiku’s failures lean the other way, with asks that go weak or a rationale that overreaches. Luna held the middle: warm and firm.

3. The winner never made things up.
The rubric’s biggest penalty was for inventing market data. Luna took zero such penalties across 30 plans; its worst miss was hedging instead of naming a firm target.

4. Cheap won on every axis, twice.
Luna was 8× cheaper than the new flash, fastest, and highest-scoring — the same sweep it pulled on our dating-bio benchmark.

The fine print

  • Judge: gpt-5.4-mini — not a contestant — scoring 0–1 blind against the rubric.
  • Sample: 30 scenarios × 3 models per round, zero failures. Run-to-run judge variance is real (±0.05–0.1 on the same model), so trust the ranking, not the second decimal; the table shows one round on the latest versions.
  • Cost: both rounds together, judging included, about $3.40.

Run this on your own use case

The flow was described in one sentence in Evaligo; each experiment variant swaps the model and nothing else — which is also why re-running on a new model version takes minutes. If your product depends on tone, this is the benchmark to run before you pick a model, and again every time a vendor ships an upgrade.

#case study#model comparison#llm judge#consumer ai#evaluation
DL

Danny Lev

Founder & CEO at Evaligo

Founder of Evaligo. Building AI automation tools that help teams ship faster. Previously led engineering at enterprise AI companies.

10+ years in AI/ML engineeringBuilt systems processing millions of AI requests

Ready to Build This?

Start building AI workflows with Evaligo's visual builder. No coding required.

✓ No credit card✓ Free tier available✓ Deploy in minutes

Need Help With Your Use Case?

Every business is different. Tell us about your specific requirements and we'll help you build the perfect workflow.

Get Help Setting This Up

Free consultation • We'll review your use case • Personalized recommendations