All Posts

Cafe Bench: GPT-6.1 Sol is incredible value, but not as good as Opus

by Anand Ani1 min read

Last week we published Cafe Bench: Dot gets the books of a four-café chain and runs it for a year, and the score is how much more profit it makes. This week OpenAI released GPT-6.1 Sol and Anthropic released Claude Sonnet 5.5, so we ran both at low, medium and high reasoning effort.

GPT-6.1 Sol at medium effort is now second on the board, behind Opus 5.5. Sonnet 5.5, on the other hand, is just bad: at low and medium effort, GPT-6 Luna beats it on both profit and cost. The Pareto frontier now runs from GPT-6 Luna to GPT-6.1 Sol (low) to Opus 5.5.

More effort did not buy more profit from GPT-6.1 Sol. This could be a fluke, but various public benchmarks show the same behavior, which is very interesting. A high-effort run of GPT-6.1 Sol also took about three and a half hours, the slowest of any model we have tested.

Sonnet 5.5, on the other hand, gets a huge boost from medium to high effort, the only setting where it made money. However, it still falls off the Pareto frontier, as it costs more than Opus 5.5 while earning less.

The full board, with this week's models in bold:

ModelYear AYear BAverageCostTime
Claude Opus 5.5 (medium)+$236,638+$216,454+$226,546$8.5722 min
GPT-6.1 Sol (medium)+$184,813+$188,084+$186,449$13.0998 min
GPT-6.1 Sol (high)+$200,084+$148,113+$174,098$27.20214 min
GPT-6 Astra (medium)+$129,942+$184,064+$157,003$56.34132 min
Claude Fable 5.1 (medium)+$75,072+$204,709+$139,890$12.8127 min
GPT-6.1 Sol (low)+$138,793+$73,347+$106,070$3.9738 min
GPT-6 Sol (medium)+$106,889+$90,213+$98,551$7.4128 min
Claude Sonnet 5.5 (high)−$17,574+$185,433+$83,929$15.6540 min
GPT-6 Astra (low)+$87,107+$63,128+$75,117$22.6851 min
GPT-5.6 Sol (medium)−$19,374−$73,116−$46,245$9.4028 min
GPT-6 Luna (medium)−$22,696−$98,103−$60,400$1.8640 min
Claude Sonnet 5.5 (medium)−$271,768+$32,737−$119,516$3.3812 min
Claude Sonnet 5.5 (low)+$4,934−$341,267−$168,167$2.179 min
GPT-5.6 Luna (medium)−$226,377−$339,861−$283,119$1.1110 min

Things to keep in mind

  • Two runs per setting. With Sonnet 5.5's years this far apart, its ranking could move a lot with more runs.
  • Same world as last week. Nothing about the simulation changed, so every row here is directly comparable with the first post.

Anand Ani

Anand is a founding AI engineer at Dot. Builder at heart, part-time nerd, and a chess player on the side.