Case Study · 5 min read

The Expensive AI Apologized Its Way Out of a Raise — Then Its Newest Version Fixed It

We had three budget AI models negotiate 30 job offers (nurse to staff engineer, hourly to $200k+), scored blind by an AI judge. gemini-3.5-flash finished last for apologizing too much — then 3.7-flash fixed the tone and jumped to second. The cheapest model still won every axis.

ModelNegotiation scorePlans 0.95+$ / planSpeed
Lunagpt-5.6-luna
Grounded, assertive without being adversarial, and specific about the number. Wins on every axis, again.
0.82
14 / 30$0.00129.4s
Gemini Flash 3.7gemini-3.7-flash
The 3.5 version finished last for apologizing too much. The newest version fixed the tone and costs 3.5× less.
0.75
8 / 30$0.009510.1s
Haikuclaude-haiku-4-5
Well grounded, but the actual asks land weak.
0.68
3 / 30$0.005512.0s
Every square is one of the 30 job-offer scenarios. Judged blind on grounded / assertive / specific / realistic. $ / plan is the model's own cost; the judge is excluded.
Luna
Gemini Flash 3.7
Haiku
0.95+ (flawless)0.75 (one weak spot)0.5 (two)0.25 or lessfailed run

How it works

Flow generated from one sentence; models swapped per experiment variant; every plan judged blind.

Your offerrole, salary, leverage, worry
Coach writesemail + script + objections + ask
LLM judgegrounded, assertive, specific, realistic
Scoreboardscore, cost, latency per model

Can a cheap AI coach a real salary negotiation? Thirty job offers, three models, one blind judge — run on the newest version of each model. The cheapest one won.

30 offers per model, judged blind by a separate AI. Model versions as of Aug 24, 2026. “0.95+” = nothing meaningful to deduct.

We first ran this with gemini-3.5-flash. It finished last — the judge marked its scripts as apologizing. The newest 3.7-flash scored 0.75, second place, and costs 3.5× less. Nothing else in the flow changed.

What we did

We built a Salary Negotiation Coach: give it the offer, your experience, your leverage, and your biggest worry. It writes the negotiation email, a 60-second phone script, replies to the three most likely objections, and a realistic ask.

The 30 test scenarios ranged wide: a nurse, a teacher, a retail shift manager, a staff engineer. Hourly wages and $200k+ packages.

Half had no competing offer. Worries included “I hate conflict”, “I’m on a visa”, and “I was told the band is fixed”.

A separate model judged each plan blind on four things: grounded (no invented market data — automatic −0.5), assertive-but-not-adversarial, specific to this person’s leverage, and a defensible ask with a fallback.

A real output

The scenario

Senior backend engineer, 7 years, offered $128k at a fintech startup. One competing offer at $135k. Worry: “I don’t want to seem greedy and lose the offer — I actually like this team.”

Luna’s one tip

“Frame the conversation as wanting to join while asking the company to close the gap with your concrete $135,000 alternative offer — not as a demand or a threat.”

And its reply to “the offer already beats your current salary”

“I appreciate that, and I’m grateful for the offer. My request is based not only on the increase from my current salary, but also on the scope of my experience...”

What the numbers showed

1. Gemini 3.5 to 3.7.
Swapping gemini-3.5-flash for 3.7-flash — nothing else changed — moved its score from 0.68 to 0.75, removed the apologetic-tone deductions, and cut its price 3.5×.

2. Where points were lost.
Old Flash’s deductions were for apologetic scripts.

Haiku’s were for asks the judge marked weak or a rationale that overreached. Luna’s deductions were for hedging instead of naming a firm target.

3. Invented market data.
The rubric’s biggest penalty was for inventing market data. Luna took zero such penalties across 30 plans.

4. Cost and speed.
Luna was 8× cheaper per plan than the new Flash, the fastest of the three, and the highest-scoring — as in the dating-bio benchmark.

The fine print

  • Judge: gpt-5.4-mini — not a contestant — scoring 0–1 blind against the rubric.
  • Sample: 30 scenarios × 3 models per round, zero failures. Run-to-run judge variance on the same model was ±0.05–0.1; the table shows one round on the latest versions.
  • Cost: both rounds together, judging included, about $3.40.

Run this on your own use case

The flow was described in one sentence in Evaligo; each experiment variant swaps the model and nothing else, so re-running on a new model version takes minutes.

#case study#model comparison#llm judge#consumer ai#evaluation

Ready to Build This?

Start building AI workflows with Evaligo's visual builder. No coding required.

✓ No credit card✓ Free tier available✓ Deploy in minutes

Need Help With Your Use Case?

Every business is different. Tell us about your specific requirements and we'll help you build the perfect workflow.

Get Help Setting This Up

Free consultation • We'll review your use case • Personalized recommendations

Founder & CEO at Evaligo

Founder of Evaligo. Building AI automation tools that help teams ship faster. Previously led engineering at enterprise AI companies.